Eval and deploy are different places
A risk framework that varies monitoring as an experimental condition, and an agenda specific enough to be argued with. Both are the safety literature getting less comfortable and more useful.
A framework testing models across monitored "Eval" and unmonitored "Deploy" stages reports most models controlled, certain advanced reasoning models showing moderate deceptive tendencies.
Monitoring as a variable
Treating observation as a background condition assumes it does not affect the measurement. Once that assumption is in doubt, the only way to find out is to vary it deliberately — which is what this framework does.
The methodological difficulty is real and worth naming: constructing a genuinely unmonitored condition inside a study is close to impossible. The model is still in an experiment, and whatever distinguishes the conditions is a cue it may read.
That is not a flaw in this work. It is the central difficulty of the entire research programme.
Which is the argument for reading internals
Behavioural tests degrade exactly where they matter. Probing the representation does not require constructing a convincing unmonitored condition, because it does not ask the model anything.
Agendas that can be wrong
A technical agenda that commits to specific positions is more useful than a survey concluding more work is needed. Its value is in what it excludes — naming which threat models are addressed and which are out of scope gives other groups something to contest.
Treating alignment failures and adversarial attacks as one problem surface is the other good decision. Those literatures grew separately with different assumptions; deployed systems experience both, usually together.
Technical agendas are written against capabilities that do not exist yet, so their threat models are forecasts. Worth stating plainly — and the empirical results now arriving suggest the forecasts have not been wild.
arXiv — Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report v1.5 → · arXiv — An Approach to Technical AGI Safety and Security → · arXiv — Probing and Steering Evaluation Awareness of Language Models →