// news · interpretability2026-08-01source: arxiv / iclr

'Size doesn't matter': cosine-scored sparse autoencoders challenge the scale-up path in interpretability

A 2026 paper, Cosine-Scored Sparse Autoencoders, argues that the quality of interpretable features does not require ever-larger dictionaries. By changing how features are scored rather than how many there are, the work pushes back on the assumption that better mechanistic interpretability means bigger, more expensive SAEs.

The claim cuts against the field's default trajectory. Sparse autoencoders became the workhorse of mechanistic interpretability by decomposing model activations into many interpretable directions, and the reflex has been to scale the dictionary — more features, more compute. Cosine scoring says the scoring function, not the feature count, is where a lot of the quality lives.

If that holds, the economics of interpretability change. SAEs are expensive to train at frontier scale, and a result that gets comparable feature quality from a smaller, cheaper dictionary makes the technique viable for more labs and more models — which matters precisely because alignment work is leaning harder on looking inside models as behavioural evaluation loses trust.

It also fits a wider 2026 pattern of interpretability turning practical: neural-operator SAEs that treat concepts as functions, domain-specific SAEs for medical text, and benchmarks like MIB to measure the field. The common thread is a move from 'can we find features at all' to 'can we find them reliably and affordably' — the sign of a subfield maturing.

See our analysis →

arXiv — Size doesn't matter: cosine-scored sparse autoencoders → · arXiv — Mechanistic interpretability with sparse autoencoder neural operators →