// blog · analysis · interpretability2026-07-31source: acl / arxiv

If two sparse autoencoders disagree, at least one is describing the method — interpretability's reproducibility problem

Interpretability's implicit promise is that a model has features and a good enough method recovers them. New work finds that features learned by sparse autoencoders differ substantially between training runs on the same activations. That is a problem about the method, not the model.

ACL 2026 work shows sparse-autoencoder features are inconsistent across training runs, undercutting the aspiration that there is a canonical feature set waiting to be found. If two SAEs trained on the same activations disagree about what the features are, at least one is describing the decomposition rather than the network — and from inside, you cannot tell which.

Why this cuts deeper than a negative result

Most negative results narrow a claim. This one questions the object. Interpretability has proceeded as though features are out there to be found, the way a biologist finds a cell type. Run-to-run inconsistency is compatible with a weaker story: that SAEs impose a decomposition which is useful without being unique, and that the interpretations we read off it are partly artefacts of the training seed.

Useful-but-not-unique is a perfectly respectable position. It is just not the position most interpretability writing has been taking.

The discipline this implies

A feature found once is a hypothesis. A feature found consistently across independent runs is a result. That is a higher bar than the field has been clearing, and adopting it will slow publication — which is the correct trade for a research programme whose entire value proposition is trustworthiness.

Consolidation, at a useful moment

A July survey pulls circuits, sparse features and symbolic reasoning into a shared frame. A field that treats sparse features as one strand among several is better placed to absorb a result that undercuts sparse features specifically. Three vocabularies for possibly-overlapping phenomena was always going to be expensive; it is cheaper to consolidate before a strand wobbles than after.

ACL Anthology — Mechanistic Interpretability Should Prioritize Feature Consistency in Sparse Autoencoders → · arXiv — Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning →