// news · interpretability2026-08-14source: arXiv

Mechanistic interpretability publishes its own list of open problems

An agenda paper setting out what the field cannot yet do — and a shift in practice from studying toy models to automated tooling capable of analysing production-scale systems. Both are signs of a discipline moving from demonstration to method.

A field that publishes its open problems is telling you it has enough structure to know what is missing. Mechanistic interpretability spent several years producing striking individual results — a circuit here, a feature there — without an agreed account of what a complete explanation would even look like.

The scale shift is the practical half. Analysis has moved from toy models chosen for tractability to automated tooling pointed at production systems, which changes the character of the claims. A hand-traced circuit in a two-layer model is an existence proof. An automated method that runs across a deployed model is a measurement, with all the error bars that implies.

That shift is what makes the open-problems exercise necessary. Once results are cited in safety cases and increasingly in regulatory documentation, "we found an interesting feature" has to become "here is what this method can and cannot establish". The second sentence is the one auditors need.

It is consistent with the field's other recent turn — concept-level explanation methods moving into a formal alignment track. Neither is glamorous. Both are what a discipline does when its outputs start being load-bearing.

See our analysis →

arXiv — Open Problems in Mechanistic Interpretability → · Zylos Research — AI Safety, Alignment, and Interpretability in 2026 →