Interpretability made the breakthrough list. That is not the same as working
MIT Technology Review named mechanistic interpretability one of its 10 Breakthrough Technologies for 2026. A consensus paper from 29 researchers across 18 organisations set out the field's open problems — which is the more useful document.
Mechanistic interpretability was named one of MIT Technology Review's 10 Breakthrough Technologies for 2026. The field has gone from a niche pursuit to a headline in about four years, and the recognition is deserved.
The more informative artefact is the other one: a paper by 29 researchers across 18 organisations establishing consensus on the field's open problems — what the key questions are, which methods work, and where the field needs to go. Twenty-nine authors from eighteen organisations agreeing on a problem list is a real milestone. It is also, read plainly, a list of things that do not yet work.
What has genuinely changed is the shift from theorising to mapping. Labs stopped speculating about what might be inside a network and started tracing sequences of features and the path from prompt to response. That is the difference between a philosophy of mind and an anatomy.
Anatomy is not yet diagnosis. Being able to trace a feature is not the same as being able to say, before deployment, that a model will not do a specific thing — and that is the claim safety cases need. The gap between "we can see inside" and "we can make guarantees" is the whole remaining problem.
Which is the honest way to hold both facts. Real progress, real recognition, and a headline finding this year that the reasoning we thought we could read was never the reasoning.
MIT Technology Review — Mechanistic interpretability: 10 Breakthrough Technologies 2026 → · IntuitionLabs — Understanding Mechanistic Interpretability in AI Models →