'Size doesn't matter' — interpretability turns from scaling up to sharpening, and gets cheaper
Every young technique has a phase where progress means bigger. Interpretability's sparse autoencoders were in it. A 2026 result arguing the scoring function matters more than the feature count marks the turn from scaling to sharpening — and it arrives exactly when alignment needs interpretability it can afford.
Cosine-scored sparse autoencoders argue that better interpretable features do not require ever-larger dictionaries — that the scoring function, not the feature count, is where much of the quality lives. If it holds, the economics of interpretability change.
Why the economics matter now
SAEs are expensive to train at frontier scale, and a result that gets comparable quality from a smaller, cheaper dictionary makes the technique viable for more labs and more models. That matters precisely because alignment is leaning harder on looking inside models as behavioural evaluation loses trust — interpretability that only the largest labs can afford cannot bear that weight.
It fits a broader turn toward rigour. The MIB benchmark and a wave of critical papers are pushing interpretability toward reproducibility — asking whether findings hold across models and seeds, and whether features should be explained from weights rather than the examples they happen to light up on.
From 'can we find features' to 'can we trust them'
The first wave of mechanistic interpretability asked whether interpretable structure existed at all. This wave asks whether the results are reproducible, affordable, and grounded — the questions you only reach once the first is settled. That is what a subfield maturing looks like, and it is the foundation any real reliance on interpretability requires.
Bigger was the easy answer. Cheaper and more reliable is the useful one.
arXiv — Size doesn't matter: cosine-scored sparse autoencoders → · arXiv — MIB: a mechanistic interpretability benchmark →