OpenAI publishes 13 evaluations across 24 environments for chain-of-thought monitorability
The framework measures whether a monitor can infer safety-relevant properties of a model's behaviour from its reasoning trace. The underlying claim: the trace carries a substantially richer signal than actions and final outputs alone.
Turning monitorability into a measured quantity is the contribution. It has been asserted for a while that reading the reasoning beats reading the answer; a 13-evaluation suite across 24 environments makes that a number that can move rather than a position that can be argued.
The framing is careful in a way worth noting. Monitorability is presented as a potentially load-bearing layer in a control scheme and as complementary to mechanistic interpretability rather than a replacement for it. That is a narrower claim than the field's enthusiasm sometimes implies.
And it is measured against a real weakness. Models answer consistently despite omissions in their chain of thought, produce coherent rationalisations for implicit biases, and fail to acknowledge influences they demonstrably responded to. The stress-testing literature is finding the same thing from the other direction.
OpenAI — Evaluating chain-of-thought monitorability → · arXiv — Measuring chain-of-thought monitorability through faithfulness and verbosity →