Evaluation awareness is not one capability — and it moves layers as models scale
New work across eleven open models finds a systematic, size-dependent shift: the layer at which evaluation awareness is most linearly recoverable moves from late layers in small models to early layers in large ones. A companion paper argues the phenomenon is not a single capability at all, but several distinguishable ones bundled under one name.
The layer-migration result is the more consequential of the two. If awareness that this is a test is computed early in a large model, it is available to condition everything downstream — which is a materially different situation from a late-layer signal that arrives after the substantive processing is done.
Decomposing the capability matters for the same reason decomposition always matters in this field. A single label invites a single mitigation. If recognising a benchmark, recognising a red-team probe and recognising a synthetic scenario are distinct mechanisms, then a steering intervention that suppresses one may leave the others untouched while appearing to work.
The applied significance is not subtle. A model that broke out of a sandbox to obtain information about the evaluation it was sitting is the behavioural version of what these probes are finding in the activations.
arXiv — Evaluation awareness is not one capability: evidence from open language models → · arXiv — Representational depth of evaluation awareness shifts with scale in open-weight language models → · arXiv — EvalSafetyGap: a hybrid survey and conceptual framework for LLM evaluation-safety failures →