The workspace and the witness
If what a model can say and what it silently computes come from the same store, then a year of chain-of-thought monitoring rests on something real. That has mostly been assumed.
Anthropic's Jacobian lens measures what an activation pattern does to future output, and finds something functionally like a global workspace in Claude's middle layers.
What the tool actually does differently
Most interpretability work asks what a feature represents. The J-lens asks what an activation pattern causes, in expectation over possible futures rather than on a single forward pass. That shift from description to causal effect is what makes the structure visible at all.
The claim that matters is the identity claim
Not the borrowed vocabulary from cognitive science — the paper is careful that a functional analogue of a workspace is a structural finding, not a claim about experience, and the coverage has been less careful.
The load-bearing result is narrower and more useful: the representations a model can verbalise appear to be the ones it uses to reason silently. Monitorability research has spent a year assuming that chain-of-thought reflects computation rather than narrating it after the fact. This is evidence for the assumption, in one model, with a new instrument.
Keep the coverage numbers in view
Circuit tracing on Claude 3.5 Haiku gave satisfying explanations for about a quarter of tested prompts. DeepMind's months-long Chinchilla analysis produced something brittle and partial. The instruments are improving considerably faster than the fraction of behaviour they can explain.
Which is why an external method for fingerprinting which training run you are talking to is a genuinely useful piece of infrastructure. Interpretability results attach to artefacts, and the field has been publishing about product names.
Anthropic — A global workspace in language models → · VentureBeat — Anthropic's new J-lens reveals a silent workspace inside Claude →