DeepMind reports tools that localise model behaviours to individual circuits
Alongside its retreat from sparse autoencoders, DeepMind reports tooling that pins specific behaviours to individual circuits — and demonstrates 'patching' alignment properties by transferring safety behaviours without full retraining. Localisation plus transfer is a more practical combination than either alone.
Localisation is the prerequisite for intervention. Knowing that a behaviour lives in an identifiable circuit means it can be measured, monitored, and — the harder claim — modified without disturbing everything else. That is a different proposition from producing an interpretable description after the fact.
Alignment patching is the consequential half. Transferring a safety property into a model without a full retrain changes the economics of fixing problems: today, discovering an alignment flaw after release often means waiting for the next training run. A patch mechanism turns that into something closer to a software update.
It also complicates the tidy story about DeepMind abandoning interpretability. Stepping back from sparse autoencoders after negative downstream results, while advancing circuit localisation and behaviour transfer, is not retreat — it is reallocation toward the techniques that survived contact with real tasks.
Google DeepMind alignment team — A summary of recent work (July 2026) → · Claude 5 Hub — AI safety 2026: alignment research breakthroughs →