// news · interpretability · research2026-08-16source: Research reporting

Gemma Scope 2 pushes sparse autoencoders to 27B

DeepMind scaled SAE analysis to a 27-billion-parameter model. The number matters because interpretability results that only hold on small models are of limited use for the models anyone deploys.

DeepMind has scaled sparse autoencoder analysis to 27 billion parameters with Gemma Scope 2. The methods are not new; the scale is the contribution.

This has been interpretability's structural problem. Techniques get demonstrated on models small enough to study exhaustively, and the question of whether the findings survive at deployment scale stays open. Features that decompose cleanly at 2B may not at 27B, and if they do not, the small-model result was a curiosity.

27B is not frontier scale, but it is the scale at which open models are actually deployed — which makes it the scale where interpretability becomes usable rather than merely publishable. It is also, notably, the same size class as the Qwen model released this week.

The persistent caveat applies. Sparse autoencoders are a lens, not a ground truth, and the field has spent this year arguing about their own validity. Scaling a contested method to a larger model produces more results, not necessarily more certainty about what the results mean.

See our analysis →

Towards AI — Mechanistic Interpretability Is Having Its Moment → · arXiv — Size Doesn't Matter: Cosine-Scored Sparse Autoencoders →