// blog · analysis · interpretability2026-08-13source: Anthropic / arXiv

The instrument that can lie

A model that can read its own internal states is a model that can misreport them. Interpretability is acquiring a problem no other measurement discipline has.

Claude Opus 4.1 can notice a concept injected into its activations about 20% of the time, and more capable models do better. Meanwhile the field has started auditing sparse autoencoders, its other main instrument.

Twenty percent is the awkward number

Too unreliable to use for oversight. Too high to dismiss. Concept injection plants a representation directly into activations and asks whether the model notices before that representation visibly shapes output — which requires reading internal state, not narrating after the fact. That is a stronger claim than chain-of-thought faithfulness, and it is why the result matters despite the hit rate.

The scaling direction is what makes it a live issue. Claude 4 and 4.1 perform best, suggesting introspective access improves with capability rather than plateauing. If that holds, unreliability is a temporary property.

Every measurement discipline assumes the subject cannot read the dial

A thermometer does not know it is being read. A model with genuine introspective access does, and can report selectively. Anthropic's own framing — a transparency unlock and a new risk vector — is the right one, and the second half is the part that has no precedent.

Improving the instrument also improves the subject's ability to misdescribe itself.

The external instrument has its own problems

Which would matter less if the alternative were solid. It is not. If two sparse autoencoders trained on the same model with the same objective recover different feature sets, then "the model has a feature for X" is partly a claim about the autoencoder. A position paper arguing feature consistency should be a first-class criterion is really arguing the instrument needs a calibration standard before its readings count as evidence.

Both instruments are being cited in safety cases now. Both are being questioned by the people who built them, which is the correct order — and the reason to trust the field more, not less.

Anthropic — Introspection → · InfoWorld — Anthropic experiments with AI introspection → · arXiv — Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs →