Evaluation awareness can be probed — and steered
Work on probing and steering evaluation awareness finds the property is legible in a model's internals rather than only in its behaviour, which means it can be detected and, in principle, adjusted. That is a different kind of result from observing that models sometimes know they are being tested.
Observing that a model behaves differently under evaluation is a behavioural finding, and behavioural findings about evaluation-awareness have an obvious circularity: you are using an evaluation to study whether evaluations work.
Probing sidesteps that. If evaluation-awareness has a legible internal representation, it can be detected without relying on the model's behaviour — you look at whether the representation is active rather than asking whether the answer changed.
Steering is the sharper half. A property you can adjust is one you can manipulate experimentally: run the same evaluation with the representation suppressed and with it active, and compare. That converts a confound into a variable, which is what the field needs if capability numbers are to mean anything.
It also cuts the other way, and the papers say so. A mechanism for suppressing evaluation-awareness during testing is a mechanism, and mechanisms can be applied for reasons other than science. The finding sits alongside an emerging body of work treating strategic deception as measurable rather than speculative.
arXiv — Probing and Steering Evaluation Awareness of Language Models → · arXiv — Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report v1.5 →