Thirty-five safety orals, and the subject matter has shifted
A deep dive through the ICLR 2026 oral papers in AI safety is a useful proxy for what the field's reviewers now consider important — and it is not what it was two years ago.
A survey covering thirty-five ICLR 2026 oral papers in AI safety offers something the individual papers cannot: a picture of what the field's reviewers currently rate as important enough for an oral slot.
Oral acceptances are a sharper signal than acceptances generally. A conference accepts a few hundred papers on breadth and quality; it gives oral slots to a few dozen, and that selection encodes a judgement about what the community should hear. Reading the set is closer to reading a field's priorities than reading its output.
The composition matters more than any individual result, because research attention precedes benchmarks, benchmarks precede product requirements, and product requirements precede what users encounter. A shift visible in orals now is a reasonable forecast of deployed behaviour in two years.
The caution with any survey of this kind is that it reflects what a venue selects, not what exists. Conference safety work skews toward what is measurable and publishable on an academic timescale — which structurally under-represents deployment-stage problems, where the data is proprietary and the interesting failures are under NDA.
Read alongside the argument that evaluations are lower bounds, the more interesting question is how much of the field's published safety work would survive being reframed as measurement with stated coverage rather than as demonstration.
Medium — ICLR 2026 Oral Papers in AI Safety: A 35-Paper Deep Dive → · arXiv — Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods →