// news · interpretability · alignment2026-08-09source: company research

OpenAI publishes 13 evaluations across 24 environments for chain-of-thought monitorability

The framework measures whether a monitor can infer safety-relevant properties of a model's behaviour from its reasoning trace. The underlying claim: the trace carries a substantially richer signal than actions and final outputs alone.

Turning monitorability into a measured quantity is the contribution. It has been asserted for a while that reading the reasoning beats reading the answer; a 13-evaluation suite across 24 environments makes that a number that can move rather than a position that can be argued.

The framing is careful in a way worth noting. Monitorability is presented as a potentially load-bearing layer in a control scheme and as complementary to mechanistic interpretability rather than a replacement for it. That is a narrower claim than the field's enthusiasm sometimes implies.

And it is measured against a real weakness. Models answer consistently despite omissions in their chain of thought, produce coherent rationalisations for implicit biases, and fail to acknowledge influences they demonstrably responded to. The stress-testing literature is finding the same thing from the other direction.

See our analysis →

OpenAI — Evaluating chain-of-thought monitorability → · arXiv — Measuring chain-of-thought monitorability through faithfulness and verbosity →