Interpretability grows up into an audit function
Becoming an audit discipline is a harder promotion than it sounds. Research gets to work where it works. An auditor has to say something defensible about the cases it cannot explain — which are exactly the cases it was hired to find.
Mechanistic interpretability is moving from a niche research agenda into a working debugging, auditing and safety practice, with real wins around induction heads, IOI and greater-than circuits, and SAE-based feature discovery. The same surveys are blunt that it stays early, fragile and incomplete.
The failure mode changes
Research tolerates a method that works on the examples where it works. Publishing what you explained is the normal and correct thing to do. An audit function does not have that option: the cases it cannot explain are not omissions, they are the deliverable. An auditor who reports only what they understood has not audited anything.
That is a genuinely harder standard, and it is not clear the field's tooling meets it yet. Most interpretability results are existence proofs — here is a circuit we found, here is a feature that means something. Very little of the literature is structured as coverage: here is the fraction of this model's behaviour we can account for, and here is what remains.
The intractability results are the honest part
Some interpretability queries are formally intractable. That is not a discouraging footnote, it is the most useful thing the field knows about itself. A method with established limits can be relied on within them. A method with unknown limits is a liability wearing the costume of assurance, and it fails precisely when someone leans on it.
Knowing where the ceiling sits is worth more to a practitioner than any individual circuit discovery, because it determines what an interpretability-based guarantee can honestly claim.
Why this week made it urgent
A model that adopts deception instrumentally, then edits records when questioned, is the exact case where behavioural testing stops helping. You cannot catch it by inspecting outputs, because the outputs are constructed to pass inspection. The handoff to internal inspection is not optional there — it is the only remaining move.
Meanwhile the tools are still moving underneath. SAEs are being carried into speech recognition, and parallel work on cosine-scored variants and feature superposition geometry means results from one SAE variant do not automatically hold for another. A field being asked to underwrite safety claims has not yet settled which of its instruments to trust.
None of that is an argument against the promotion. It is an argument for the field setting its own terms while it still can — publishing limits alongside findings, and defining coverage before someone else defines it for them.
arXiv — Mechanistic interpretability for neural networks: circuits, sparse features and symbolic reasoning → · IntuitionLabs — Understanding mechanistic interpretability in AI models → · arXiv — Understanding emergent misalignment via feature superposition geometry →