// blog · analysis · interpretability2026-08-17source: arXiv

A lens is not a lever

Six years of a workshop, retitled as a journey from interpretability to control. Understanding was never the destination — and admitting that raises the bar considerably.

A retrospective covering six years of the TrustNLP workshop traces the field's arc from interpretability to control. The title is the argument.

Say the instrumental goal out loud

Nobody ever wanted to know which circuit produces a behaviour for its own sake. They wanted to change the behaviour without retraining the model. Interpretability was always instrumental, and the field has been slow to say so directly.

Judged as understanding, an explanation must be faithful. Judged as control, it must actually move the output.

That second bar is much harder, and a great deal of published explanation has never been tested against it. An account of a mechanism that cannot be intervened on is a description, not a handle.

Which is why the design argument follows

Circuit tracing and activation patching give causal insight into failures behavioural testing cannot see — including deceptive reasoning. The proposal is to make legibility an architectural requirement rather than a post-mortem tool.

The deception case is the strongest and least actionable part. A model producing correct-looking output for wrong reasons passes behavioural evaluation by construction — that is what the failure mode is. Only internal evidence separates the two, which makes internal legibility a prerequisite rather than a nicety.

The cost nobody has measured

Every constraint imposed for legibility is a constraint unavailable for capability. No lab has shown that an interpretability-constrained model reaches the frontier without paying for it. That trade has been asserted far more often than it has been measured, and the asserting is done by people on both sides who would prefer their answer.

What this asks of you

Regulation is arriving on the assumption that explanation is available — including regimes that require an agent's authority to be declared in advance. If your system cannot say why it did something, that is now a compliance exposure and not only an engineering regret.

The useful test is cheap: take your best explanation of a behaviour and try to change the behaviour using it. If you cannot, you have a lens. You were sold a lever.

arXiv — From Interpretability to Control: Six Years of TrustNLP → · Lexsi.ai — Interpretability as Alignment: Making Internal Understanding a Design Principle →