// news · interpretability2026-08-04source: aisocratic / anthropic

Anthropic's turn-averaged autoencoders compress millions of activations into per-turn behavior signals

Anthropic's latest circuits update introduces turn-averaged sparse autoencoders — collapsing millions of per-token activations into a handful of per-conversation-turn features. It is interpretability re-engineered for the question production actually asks: not what the model computed on each token, but what it is doing this turn.

The unit of analysis is the innovation. Token-level features are the natural resolution for understanding computation and a hopeless resolution for monitoring behavior — no operator can watch millions of activations stream past. Averaging over a conversational turn produces features at the resolution operators think in: is the model being evasive, sycophantic, deceptive, off-task, right now, in this exchange.

This is interpretability converging on its deployment form: a monitoring layer. The microscope work — circuits for addition, emotion vectors that causally steer behavior — established that internal structure is real and readable. Turn-averaged features are that readability packaged into something a safety pipeline can consume in real time, which is what it takes for interpretability to become a standard component rather than a research demo.

The contrast with DeepMind's simultaneous retreat from SAEs sharpens the moment. One lab found autoencoders failing on downstream tasks and moved to probes; the other is betting the fix is changing what the autoencoder summarizes. The field's two leading teams now disagree in public about its central technique — and behavioral monitoring at scale is the arena where the disagreement will be settled.

See our analysis →

AI Socratic — AI Socratic July 2026 — lost in J-space → · Anthropic — Interpretability research →