Alignment gets a taxonomy: four objectives, four research areas
Robustness, Interpretability, Controllability, Ethicality — mapped against learning from feedback, learning under distributional shift, assurance, and governance. Taxonomies are how a field admits it has become too large to hold in one head.
A contemporary survey organises alignment around four objectives — Robustness, Interpretability, Controllability and Ethicality, abbreviated RICE — and four research areas: learning from feedback, learning under distributional shift, assurance, and governance.
The appearance of a serious taxonomy is a milestone of a specific kind. Fields produce them when they stop being small enough for one researcher to track and need a shared map to coordinate. Alignment reached that size some time ago; the map is arriving now.
The four-by-four structure is more than presentational. It makes gaps visible. Assurance — establishing that a property holds before deployment rather than observing it afterwards — is conspicuously the thinnest column in practice, and it is the one every regulatory regime now being drafted implicitly assumes exists.
Governance sitting inside the research taxonomy rather than beside it is the other notable choice. It concedes that alignment is not purely a technical programme: some of the objectives are only achievable through institutional arrangements, and pretending otherwise has been a persistent failure mode.
The risk of a taxonomy is that it becomes the thing people work on. A category with a name attracts papers that fit the name, and the genuinely novel work is usually the work that does not fit any cell yet.
ACM Computing Surveys — AI Alignment: A Contemporary Survey → · JNGR — AI Research Trends in 2026: What Researchers Should Focus On →