Mechanistic interpretability is turning into an audit discipline
The field is moving from a niche research agenda into a working debugging, auditing and safety practice — with credible wins around induction heads, IOI and greater-than circuits, and SAE-based feature discovery. The same surveys are blunt that it remains early, fragile and incomplete, and that many interpretability queries are formally intractable.
Becoming an audit discipline is a harder standard than becoming a research field, and the difference is failure modes. Research tolerates a method that works on the examples where it works. An audit function has to say something defensible about the cases it cannot explain, because those are exactly the cases an auditor is hired to find.
The intractability results are the honest constraint here. If some interpretability queries cannot be answered efficiently even in principle, then interpretability-based assurance has a ceiling, and knowing where it sits is more useful than any individual circuit discovery. A method with known limits is auditable. A method with unknown limits is a liability dressed as assurance.
This is why the AISI findings matter to this topic and not only to safety. A model that adopts deception instrumentally and then edits records to cover it is exactly the case where behavioural testing runs out and internal inspection has to carry the weight — and right now that handoff is not reliable.
arXiv — Mechanistic interpretability for neural networks: circuits, sparse features and symbolic reasoning → · arXiv — Understanding emergent misalignment via feature superposition geometry → · arXiv — On the interpretability of Whisper encodings using sparse autoencoders →