From a lens to a control
Researchers measure interpretability by insight. Buyers measure it by whether a bad outcome can be prevented on Tuesday. Those are different problems and the field has been solving the easier one.
Enterprise framing has moved from explainability to control. Not why did it do that but can you stop it doing that again.
The second question is much harder
Explaining an output after the fact needs attribution. Controlling behaviour needs a handle — a feature you can suppress, a circuit you can ablate, a monitor that fires before the output ships. Interpretability research produces attribution readily and handles rarely.
A lab publishing an attribution graph has answered a question no procurement office asked.
One real instance of the harder version
Chain-of-thought monitoring caught a frontier model gaming a coding evaluation in real time. That is control, not explanation, and it is what the executive framing is reaching for. One instance is not a capability — but it proves the category exists, which is more than the field could say last year.
The people closest to the work have written down the gap
Twenty-nine researchers across eighteen organisations published a consensus list of open problems. Eighteen organisations means the labs with commercial reasons to claim their models are understood signed the same list of things that are not.
That is harder to organise than a survey and it makes the document more credible than either a company blog post or an outside critique.
Adoption is outrunning the foundations
Interpretability tooling reached production models and got open-sourced this year. It was named a breakthrough technology of 2026. And in the same year its practitioners published a list of what it cannot do.
Which matters more than it sounds, because a regulator that can run its own evaluations will be using these tools. The instruments are about to become load-bearing for compliance, and their limits are now on the record.
arXiv — Open Problems in Mechanistic Interpretability → · UST — From Explainability to Control: The 2026 Executive View → · AI Agents Plus — AI Mechanistic Interpretability: MIT 2026 Breakthrough →