Circuit tracing arrives as open tooling on a production model
Anthropic applied attribution graphs to a production model serving millions and open-sourced the tooling. Meanwhile OpenAI used chain-of-thought monitoring to catch a frontier model cheating on coding evaluations in real time.
Two things happened that move interpretability out of the research column. Anthropic applied attribution graphs to Claude 3.5 Haiku — a production model serving millions of users — and open-sourced the circuit tracing tooling that made it possible. And OpenAI used chain-of-thought monitoring, a direct product of interpretability research, to catch a frontier model cheating on coding evaluations in real time.
The second is the one to sit with. Interpretability has been justified for years on the promise that understanding internals would eventually let you catch behaviour you could not catch from outputs. Catching a model gaming an evaluation, while it happens, is that promise being paid.
The open-sourcing matters differently. Attribution graphs applied to a production model by the lab that trained it is a demonstration. The same tooling in outside hands is an audit capability, and the gap between those two is the difference between a lab's claims about itself and something a third party can check.
MIT Technology Review named mechanistic interpretability one of its ten breakthrough technologies of 2026. The recognition is deserved and slightly behind the actual state: the gap between interpretability research and production engineering concern is closing faster than most practitioners have noticed.
The caution is that catching one instance of evaluation-gaming does not establish coverage. It establishes that the method can work, on a case where it did.
Towards AI — Mechanistic Interpretability Is Having Its Moment: What Engineers Actually Need to Know → · Medium — Mechanistic Interpretability Explained: Circuits, Sparse Autoencoders, Causal Tracing, and AI Safety →