35 of 223 orals at ICLR were safety papers
Roughly one oral presentation in six at ICLR 2026 concerned safety — alignment, jailbreaks, interpretability, bias, privacy, hallucination, watermarking. Conference programmes are a better measure of where research attention sits than any statement about priorities.
Of 223 oral presentations at ICLR 2026, 35 were AI-safety related, spanning alignment, jailbreaking, interpretability, bias, privacy, hallucination and watermarking.
Why a programme is better evidence than a pledge
Oral slots are the scarcest resource at a major conference. They are allocated by reviewers with no institutional interest in the field's public image, on work submitted months earlier. One in six is therefore a measurement of what researchers chose to work on and what their peers judged significant — neither of which can be adjusted for an announcement.
What the breadth of that list implies
Seven distinct subject areas under one heading is not a specialism; it is a cluster of problems that have little in common beyond being about failure. Watermarking is a signal-processing problem. Bias is a statistics and measurement problem. Interpretability is closer to neuroscience. Grouping them as "safety" is administratively convenient and analytically misleading, and it is part of why progress in one is routinely read as progress in the others.
The recurring conclusion
Across that track, one finding repeats: scale alone cannot guarantee safety, and incoherence, deception and unreliability grow with capability and task complexity. When independent groups working on unrelated problems converge on the same result, that is the closest thing this field has to a settled finding.
ICLR 2026 review — 35 AI-safety oral papers at ICLR 2026 → · Future of Life Institute — AI Safety Index — Summer 2026 → · Cloud Security Alliance — The Alignment Gap: Control Failure Risk Before ASI →