Two models argue, a smaller one judges, and it agrees with the experts 95% of the time
A deployed hybrid system has two models take opposing positions on safety-critical decisions while a smaller judge model evaluates the debate. Reported agreement with human expert panels is 95% on complex scenarios.
A hybrid arrangement reported in DeepMind's 2026 work has two models argue opposing viewpoints on safety-critical decisions, with a smaller judge model evaluating the exchange. Agreement with human expert panels on complex scenarios is reported at 95%.
State the provenance plainly: this comes to us through a survey of 2026 alignment progress rather than a paper we could open, so the 95% is a reported figure and not one with a methodology attached in front of us. It is worth reporting because the architecture is checkable even where the number is not.
Debate as a safety mechanism is an old proposal with a specific appeal: it does not require the judge to be smarter than the debaters. A weaker evaluator can adjudicate between two stronger arguing parties, because the work of finding the flaw is done by the opponent rather than by the judge. That is the property that makes it scale in principle.
The property that makes it fragile in practice is collusion. Two models from the same family, trained on the same data, may share the blind spot that the debate was supposed to surface — and a judge that agrees with experts 95% of the time is by construction wrong on the 5% where the hard cases live.
It sits in a month where the field has been unusually direct about its own limits. Thirty nations' nominated experts concluded that alignment research has not kept pace with capability. A working mechanism that needs no superhuman judge is exactly the sort of thing that gap requires.
Claude 5 Hub — AI Safety 2026: Alignment Progress and Open Challenges → · Zylos Research — AI Safety, Alignment, and Interpretability in 2026 →