// blog · analysis · interpretability2026-08-08source: arXiv preprints and company disclosure

The model knows more than it says

Influential signals show up in the reasoning trace 87.5 percent of the time and in the answer 28.6 percent of the time. Two thirds of the model's own account of itself never reaches the user.

A study of chain-of-thought under knowledge conflict tests introspective faithfulness — whether stated reasoning reflects the certainty state that actually drove the decision. Thinking tokens acknowledge the influential signal 87.5 percent of the time. Stated output: 28.6 percent.

What that gap costs

Any monitoring regime reading only the final answer is discarding most of what the model recorded about its own decision. Any user judging trustworthiness from the answer is doing the same, with less awareness that they are doing it.

The optimistic reading is that the information exists and can be read at the right layer. The uncomfortable reading is that a systematic three-to-one suppression between reasoning and output is a behaviour nobody designed, nobody requested, and nobody currently explains.

Monitorability is two things

Separating faithfulness from verbosity is the cleaner framing. A trace can be perfectly honest and too terse to catch anything. It can be exhaustive and unfaithful. Monitoring needs both, and most evaluations measure neither cleanly.

Worse, the measurement is unstable. Work on classifier sensitivity finds faithfulness scores depend materially on how faithfulness is operationalised — which makes cross-paper comparison of these numbers weaker than the numbers look.

This is not academic this week

Astra's safeguards have monitors reading the chain of thought and interrupting high-risk activity mid-run. That control is exactly as good as trace monitorability is.

A safety mechanism deployed on a Critical-rated cyber model depends on a property the research community cannot yet reliably measure. That is not an argument against the mechanism — it is the best one available. It is an argument for funding the measurement at the same urgency as the capability.

arXiv — Do models know why they changed their mind? Interpretability and faithfulness of chain-of-thought under knowledge conflict → · arXiv — Measuring chain-of-thought monitorability through faithfulness and verbosity → · OpenAI — Responding to the next frontier of critical cyber capabilities →