Truthfulness went from absent to 37% of alignment papers
A survey of the literature finds truthfulness is the fastest-growing dimension — missing entirely in 2021–2022, more than a third of papers by 2025–2026. What a field measures is what it eventually optimises.
A contemporary survey of alignment research finds truthfulness is the fastest-growing dimension in the literature: absent in 2021–2022, comprising 37% of papers by 2025–2026. Fairness remains the most consistent theme across the period, and explainability shows a U-shaped trajectory, resurging through mechanistic interpretability.
The composition of a literature is a leading indicator of what gets built. Research attention precedes benchmarks, benchmarks precede product requirements, and product requirements precede the thing users actually experience. A third of the field working on truthfulness in 2026 is a reasonable predictor of what deployed systems will be optimised for in 2028.
The rise is easy to explain and worth stating plainly: models became fluent enough that confident wrongness became the dominant failure mode. When output was obviously broken, truthfulness was not the binding problem. When output is polished and sometimes false, it is the only problem.
The explainability U-shape is the more interesting curve. Interest fell as attention moved to capability and behavioural evaluation, then returned once it became clear that behavioural testing cannot distinguish a model that is right from one that has learned what right-looking answers look like. That is the case now being made explicitly.
Survey composition is not the same as progress. Counting papers measures where effort went, not what it achieved, and fashionable dimensions attract work that would have been done under a different label anyway.
arXiv — Mapping Trustworthiness in Large Language Models: A Bibliometric Analysis Bridging Theory to Practice → · arXiv — Understanding AI Trustworthiness: A Scoping Review of AIES & FAccT Articles →