Let them argue
A safety mechanism that does not need the judge to be smarter than the thing being judged. That property is rare enough to be worth understanding even before the number is verified.
Say the provenance first
That figure reaches us through a survey of 2026 alignment progress, not a paper with a methodology in front of us. It is worth reporting because the architecture is checkable in principle even where the number is not yet.
Why the architecture is the interesting part
Debate has a specific and unusual property: the judge does not have to be as capable as the debaters. Finding the flaw in an argument is work, and in a debate that work is done by the opponent. The judge only has to notice who did it better.
Almost every other oversight proposal requires the overseer to be at least as strong as the overseen. This one does not, and that is why it might scale.
Where it breaks
Collusion, and not the deliberate kind. Two models from the same family, trained on overlapping data, may simply share the blind spot the debate was meant to surface. Neither one raises it because neither one sees it.
And a judge that agrees with experts 95% of the time is, by construction, wrong on 5% — which is where the hard cases live. A mechanism that is excellent on the easy majority and unreliable on the difficult minority has an uncomfortable resemblance to the problem it was built for.
Why this month needed it
Thirty nations' nominated experts concluded that alignment research has not kept pace with capability. The largest training cohort to date runs at nearly one mentor per fellow — an apprenticeship ratio, which tells you the constraint is people rather than money.
Capability scales with compute, and compute can be bought. Oversight currently scales with expert attention, and it cannot. A mechanism that lets weaker judges supervise stronger systems is the only kind of answer that closes that gap rather than restating it.
Claude 5 Hub — AI Safety 2026: Alignment Progress and Open Challenges → · Zylos Research — AI Safety, Alignment, and Interpretability in 2026 → · SPAR — Spring 2026 Projects →