// news · interpretability · alignment2026-08-08source: arXiv preprint

Models acknowledge an influential signal 87.5% of the time in thinking tokens — and 28.6% in the answer

A study of chain-of-thought under knowledge conflict tests introspective faithfulness: whether stated reasoning reflects the certainty state that actually drove the decision. The gap between what appears in the reasoning trace and what survives into the output is threefold.

The two numbers are the finding. Thinking tokens acknowledge the influential signal 87.5 percent of the time. Stated-output acknowledgment drops to 28.6 percent. The information is present internally and disappears on the way to the user.

That has a direct operational consequence. Any monitoring regime that reads only the final answer is discarding roughly two thirds of the model's own account of why it decided what it decided — and any user evaluating trustworthiness from the answer alone is doing the same.

It also complicates the reassuring version of chain-of-thought monitoring. The trace is more informative than the output, which is good, and the fact that models systematically omit from the answer what they acknowledged while reasoning is a behaviour nobody designed and nobody yet explains.

See our analysis →

arXiv — Do models know why they changed their mind? Interpretability and faithfulness of chain-of-thought under knowledge conflict → · arXiv — Reasoning theater: disentangling model beliefs from chain-of-thought →