Twenty-nine researchers across eighteen organisations agree on what is unsolved
A landmark paper establishes consensus on the open problems in mechanistic interpretability. A field agreeing on its own unsolved list is a sign of maturity and, read carefully, a warning.
Twenty-nine researchers from eighteen organisations have published a consensus statement on the open problems in mechanistic interpretability.
Papers like this do something specific for a field. They convert a scatter of individual scepticisms into a shared agenda, and they make it citable. A funder, a regulator or a graduate student can now point at one document instead of reconstructing the state of the art from twenty preprints that disagree.
The signal in the authorship is worth reading. Eighteen organisations means the labs with commercial interests in claiming their models are understood signed the same list of things that are not understood. That is a harder thing to organise than a survey and it makes the resulting document more credible than either a company blog post or an outside critique.
The warning is in the timing. This lands in the same year that interpretability tooling reached production models and got open-sourced, and mechanistic interpretability was named a breakthrough technology of 2026. Adoption is arriving faster than the foundations, and the people closest to the work have now said so in writing.
arXiv — Open Problems in Mechanistic Interpretability → · Zylos Research — AI Safety, Alignment, and Interpretability in 2026 →