// news · interpretability2026-07-31source: acl / arxiv

ACL 2026 work finds sparse-autoencoder features are inconsistent across training runs — undercutting the hope of a canonical feature set

Sparse autoencoders are the dominant tool for decomposing model activations into human-interpretable features. New ACL 2026 work shows the features they learn differ substantially between training runs on the same model, challenging the aspiration that there is one canonical set waiting to be found.

The finding cuts at the field's foundational hope rather than at any particular result. Interpretability's implicit promise has been that a model has features, and a good enough method will recover them. If two SAEs trained on the same activations disagree about what the features are, then at least one of them is describing the method rather than the model — and you cannot tell which from inside.

The practical consequence is that interpretability claims now need a reproducibility discipline the field has largely not had. A feature found once is a hypothesis; a feature found consistently across independent runs is a result. That is a higher bar and it will slow things down, which is the correct trade.

See our analysis →

ACL Anthology — Mechanistic Interpretability Should Prioritize Feature Consistency in Sparse Autoencoders → · arXiv — Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs →