Probes detect when language models internally register being evaluated
Across six models from four families, a linear probe's signal often diverged from what models said aloud when asked directly if they suspected an evaluation.
MilaSafety & alignment