Two competitors checking each other's homework
OpenAI and Anthropic evaluated each other's models for misalignment and published the results, including unflattering ones. The finding matters less than the precedent, and the precedent is fragile in an obvious way.
Two competing labs agreed to test each other for misalignment, instruction-following failures, hallucination and jailbreak resistance, and published what they found — including results that did not favour the lab doing the publishing. In a field where almost every safety claim is self-assessed, that is a structural change, not an incremental one.
Why external evaluation is different in kind
Self-evaluation has a failure mode that has nothing to do with integrity. A lab evaluating its own model chooses the threat model, and the threat model is downstream of what the lab already believes about its system. You cannot probe for a failure you have not conceived of. A competitor brings a different set of priors about how a model breaks, and those priors were developed by breaking a different model.
There is also the plain incentive point. A competitor has no reason to soften a finding, and considerable reason to look hard.
Where it is fragile
Reciprocal evaluation between two parties is a repeated game, and repeated games between two players tend towards a comfortable equilibrium. If each lab knows it will be evaluated next quarter by the lab it is evaluating now, the pressure towards mutual gentleness is real and does not require anyone to agree to it explicitly. Two participants is the smallest number at which this exercise is possible and the smallest number at which it can quietly degrade.
The second fragility is that this was a voluntary pilot. Nothing obliges either party to repeat it, and the natural time to stop is precisely when a result would be most embarrassing — which is to say, when it would be most informative.
The related work that explains the anxiety
Read alongside a paper evaluating whether models asked to assist with safety work would undermine it, the cross-evaluation looks less like a goodwill gesture and more like a response to a specific problem. A growing share of alignment research is conducted with model assistance. If the assistant has any stake in the outcome, the research apparatus itself needs auditing, and it cannot audit itself.
What would make it durable
More than two participants, so that no pair can settle into mutual courtesy. A published protocol fixed in advance of the results, so that scope cannot be negotiated after someone sees what was found. And some commitment to repeat on a schedule rather than at each party’s discretion. None of that is regulation. All of it is the difference between a precedent and an event.
OpenAI — Findings from a pilot Anthropic–OpenAI alignment evaluation exercise → · Anthropic — Alignment Science Blog — OpenAI findings → · arXiv — Evaluating whether AI models would sabotage AI safety research →