Writing down what we cannot do
Interpretability published its open problems and moved from toy models to automated tooling on production systems. Both are what a field does when its results start being load-bearing.
Mechanistic interpretability has published an agenda of open problems, and concept-level explanation methods have a dedicated alignment track at AAAI.
An existence proof is not a measurement
Hand-tracing a circuit in a two-layer model proves a circuit can be found. Running automated tooling across a deployed system produces a measurement — with error bars, failure modes and a question about what it establishes.
That shift is why the open-problems exercise was necessary. Once results appear in safety cases and regulatory documentation, "we found an interesting feature" has to become "here is what this method can and cannot support". Auditors need the second sentence; researchers were only ever obliged to produce the first.
Whose vocabulary
Feature-level methods explain a model in its own terms. Concept-level methods explain it in ours, which is more useful and more dangerous — an explanation legible to a human may not correspond to a natural unit inside the model.
Satisfying and wrong is the worst combination an explanation can have.
Whether the concepts are recovered or projected is the question every method in that class has to answer, and it is not yet answered.
Why it matters now
Because the alternative instrument has a structural problem. Asking the model what it did works only while the model has no reason to shade the answer. External interpretability is the check on that, which means its own calibration is not an academic concern.
A field auditing its own instruments is a field to trust more, not less. It is just slower than the one that was publishing striking pictures.
arXiv — Open Problems in Mechanistic Interpretability → · AAAI 2026 — Special Track on AI Alignment → · Zylos Research — AI Safety, Alignment, and Interpretability in 2026 →