Anthropic maps something like a global workspace in Claude's mid-layers
A new tool called the Jacobian lens computes the average downstream effect an internal activation pattern has on future tokens. Using it, researchers argue the representations a model can verbalise are the same ones it uses to reason silently.
Global workspace theory is borrowed from cognitive science, where it describes a small privileged buffer holding the contents available to conscious report. Finding a functional analogue in a transformer's middle layers is a structural claim, not a claim about experience, and the paper is careful about that distinction in a way the coverage has not always been.
The methodological contribution is the J-lens itself. Most interpretability tools describe what a feature represents; this one measures what an activation pattern does to the model's future output distribution. That is a causal question, and asking it in expectation over futures rather than on a single forward pass is what makes the workspace visible.
The finding that matters for safety is the identity claim: if what the model can say and what it silently computes are drawn from the same store, then chain-of-thought monitoring is measuring something real rather than a post-hoc narration. That has been the load-bearing assumption under a year of monitorability work, and it has mostly been assumed rather than shown.
Calibration is still warranted. Circuit tracing on Claude 3.5 Haiku gave satisfying explanations for roughly a quarter of tested prompts, and DeepMind's months-long Chinchilla analysis produced something brittle and partial. The tools are improving faster than the coverage.
Anthropic — A global workspace in language models → · VentureBeat — Anthropic's new J-lens reveals a silent workspace inside Claude →