What the field decided to measure
Truthfulness went from absent to 37% of alignment papers in four years. Research composition is a leading indicator, and this one predicts what 2028's products optimise for.
Truthfulness is the fastest-growing dimension in alignment research: absent in 2021–2022, 37% of papers by 2025–2026. Fairness stayed the most consistent theme. Explainability traced a U.
Composition predicts products
Research attention precedes benchmarks. Benchmarks precede product requirements. Product requirements precede what a user experiences. A third of the field working on truthfulness now is a reasonable forecast of what deployed systems optimise for two years out.
Truthfulness became urgent the moment output got polished enough for wrongness to stop being obvious.
When models produced visibly broken text, truthfulness was not the binding constraint. Fluency solved, it became the only constraint.
The U is the interesting shape
Explainability fell out of fashion as attention moved to capability and behavioural evaluation, then returned through mechanistic interpretability. The reason for the return is specific: behavioural testing cannot distinguish a model that is right from one that has learned what right-looking answers look like.
That is precisely the argument now being made for treating interpretability as a design constraint rather than a diagnostic.
A taxonomy arrives
Robustness, Interpretability, Controllability, Ethicality — against learning from feedback, distributional shift, assurance and governance. Fields produce maps when they outgrow one researcher's head.
The map's value is that it shows gaps. Assurance — establishing a property before deployment rather than observing it after — is the thinnest column in practice and the one every regulatory regime being drafted assumes already exists.
What this asks of you
Counting papers measures where effort went, not what it achieved, and a named category attracts work that would have happened under another label. Read the composition as a forecast of attention, not of progress.
Then check your own roadmap against the thin column. If your assurance story is a system card, you are relying on the one instrument the field has just agreed is insufficient.
ACM Computing Surveys — AI Alignment: A Contemporary Survey → · Zylos Research — AI Safety, Alignment, and Interpretability in 2026 →