Claude can sometimes report its own internal states — about 20% of the time
Anthropic's introspection work uses concept injection to test whether a model can notice a representation planted in its activations before that representation visibly shapes its output. Claude Opus 4.1 does so in roughly 20% of trials. More capable models do better, which the researchers frame as both a transparency unlock and a new risk vector.
Twenty percent is a strange number to build on, and the researchers say so. It is far too unreliable to use as an oversight mechanism and far too high to dismiss as noise, which puts the result in the most awkward category a safety finding can occupy.
The method is what gives it force. Concept injection plants a representation directly into the model's activations — the idea of "all caps", or a prefilled anomalous word — and then asks whether the model notices before the injected concept obviously changes what it says. Detecting it requires something reading internal state rather than reading output. That is a meaningfully stronger claim than a model describing its reasoning after the fact.
The scaling direction is the part with consequences. Claude 4 and 4.1 perform best, suggesting introspective access improves with capability rather than saturating. If that holds, the reliability problem is temporary.
Which is exactly why Anthropic flags it as a risk vector as well as a tool. A model with genuine access to its own internal states is a model that can report them selectively. Every gain in introspective accuracy is also a gain in the precision with which a model could misdescribe itself — and unlike external probes, this instrument is one the subject can read.
Anthropic — Introspection → · InfoWorld — Anthropic experiments with AI introspection → · Towards AI — Can AI Introspect? Anthropic's New Study Offers a Glimmer of Self-Awareness →