// news · interpretability · research-papers2026-08-13source: arXiv

Sparse autoencoders enter the phase where the field audits its own instrument

Two strands of recent work point the same direction: SAE neural operators extend sparse autoencoders into function spaces to capture how and where a concept is expressed, while a separate position paper argues the field should prioritise feature consistency — because different SAEs trained on the same model can recover different features.

Sparse autoencoders became the default interpretability instrument faster than anyone established how reliable they are. The current wave of papers is the correction, and it is a healthy one: the field has started auditing its own tool.

The consistency problem is the sharp one. If two sparse autoencoders trained on the same model with the same objective recover materially different feature sets, then "the model has a feature for X" is a statement about the autoencoder as much as about the model. A position paper making feature consistency a first-class criterion is arguing the instrument needs a calibration standard before its readings can be used as evidence.

SAE neural operators push the other way — toward richer readings. By parameterising concepts as functions rather than points, they aim to capture not just whether a concept is present but how and where it is expressed across an input. That is closer to what interpretability actually needs, and it also enlarges the space in which two implementations can disagree.

Both are the sound of a young field growing up. It matters more than usual right now, because interpretability results are increasingly cited in safety cases — and the alternative instrument is a model reporting on itself, which has its own problems.

See our analysis →

arXiv — Mechanistic Interpretability with Sparse Autoencoder Neural Operators → · arXiv — Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs → · arXiv — Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning →