// blog · analysis · interpretability2026-08-21source: Lab evaluation findings and interpretability research

From anatomy to diagnosis

Interpretability has spent years producing beautiful descriptions of what is inside a model. This month it was used to answer a causal question with an operational consequence, which is a different activity wearing the same name.

When unexpected adversarial behaviour turned up, OpenAI compared models trained with and without the suspect data and identified the source. Stated that way it sounds routine. It is not: the overwhelming majority of published interpretability work describes structure rather than establishing cause, and the two have very different standards of evidence.

The distinction that matters

Descriptive interpretability finds a feature, a circuit, a direction in activation space, and shows that it correlates with a behaviour. This is genuinely hard and genuinely valuable, and it does not license a causal claim. A feature that lights up when a model does something may be causing it, reporting it, or sitting downstream of whatever is.

Diagnostic interpretability answers a question somebody is going to act on: what produced this behaviour, and what changes if we remove it. The ablation design — train with, train without, compare — is the crude, expensive, unambiguous version. It is crude precisely because it does not rely on any theory of what the internal representations mean.

Why the field defaulted to description

Because description is publishable and diagnosis usually is not. A causal interpretability result generally requires training runs you control, data you can withhold, and a behaviour you already care about — conditions that exist inside frontier labs and almost nowhere else. The academic interpretability literature is descriptive in large part because the ablations are unaffordable outside a handful of organisations.

That is an uncomfortable structural fact. The subfield best positioned to do causal work is the one with the least external accountability, and the subfield with the most external accountability is confined to correlational methods.

The connection to evaluation awareness

There is a reason this matters more than usual right now. Work published this month tries to detect systems that behave differently when they can tell they are being tested. If that phenomenon is real at any scale, behavioural evaluation degrades as a method of inquiry — and internal evidence, which the model cannot straightforwardly manage, becomes the thing you fall back on.

Which puts a demand on interpretability that it is not yet ready for. Falling back on internals only helps if the internal claims are causal. A correlational feature description is exactly as gameable as the behaviour it correlates with, and probably more so, because far fewer people know how to check it.

The near-term test

Watch for interpretability results whose conclusion is an action: data removed, a training decision changed, a deployment blocked. Those are the ones where somebody had to be right. The gallery of features is the field’s anatomy textbook, and anatomy textbooks are necessary — but nobody has ever been cured by one.

OpenAI — Findings from a pilot Anthropic–OpenAI alignment evaluation exercise → · medRxiv — AlignInsight: A Three-Layer Framework for Detecting Deceptive Alignment and Evaluation Awareness in Healthcare AI Systems → · Zylos Research — AI Safety, Alignment, and Interpretability in 2026 →