// news · research-papers2026-08-01source: arxiv / arxiv

Sparse autoencoder neural operators parameterize concepts as functions, not scalars — capturing where and how a concept is expressed

A new paper introduces sparse autoencoder neural operators (SAE-NOs), which represent a concept as a function over the input domain rather than a single activation value. The result captures not just whether a concept is present but how and where it is expressed — a richer object than the scalar features SAEs have used to date.

The shift from scalar to function is a genuine expansion of what a feature can encode. A standard SAE feature answers a yes/no-ish question: is this concept active, and how strongly? Parameterising the concept as a function lets the representation say where in the input the concept lives and how it varies across it — spatial and structural information a scalar throws away.

The motivation is the same reproducibility pressure driving the rest of interpretability. If features are going to be the units alignment reasons about, they need to carry enough information to be checked and steered precisely, and a function-valued feature gives more surface to verify against than a single number that could match many different internal states.

It sits alongside a cluster of 2026 SAE work — cosine scoring, domain-specific training, weight-based explanation — that collectively reads as the technique growing up. The first wave asked whether SAEs could find interpretable structure at all; this wave asks what the right representation of that structure is, which is the question you only reach once the first is settled.

See our analysis →

arXiv — Mechanistic interpretability with sparse autoencoder neural operators → · arXiv — Sparse autoencoders reveal interpretable and steerable features in VLA models →