// news · alignment · safety2026-08-21source: OpenAI / Anthropic

OpenAI and Anthropic ran evaluations on each other's models

Two competitors agreed to test each other for misalignment, instruction-following failures, hallucination and jailbreak resistance, and published what they found — including results that did not favour the lab publishing them.

OpenAI and Anthropic conducted a joint safety evaluation, each running in-house misalignment evaluations against the other's public models. Both organisations published findings. The exercise is a first of its kind between labs at this level of competition.

The results are mixed in a way that is more persuasive than a clean sweep would be. Claude 4 models performed well on evaluations stress-testing respect for the instruction hierarchy, giving the best performance of any model tested on avoiding system-message conflicts. In simulated settings with some model-external safeguards disabled, OpenAI's o3 and o4-mini reasoning models were found to be aligned as well as or better than Anthropic's models overall.

Each lab therefore published a result in which it did not win. That is the part worth noticing. Self-published safety evaluations have an obvious credibility problem, and the standard fix — independent third-party auditing — has been slow to arrive because the tooling and access requirements are severe. Mutual evaluation between competitors is a cheaper approximation: neither party has an incentive to flatter the other.

The limits are real. Both labs chose the evaluations. Neither has an incentive to surface a class of problem they both share. And a pilot is not a standing regime. But it is the first arrangement in this field where a safety claim about a frontier model was checked by someone with a motive to find fault, and published anyway.

See our analysis →

OpenAI — Findings from a pilot Anthropic–OpenAI alignment evaluation exercise → · Anthropic — Alignment Science Blog — OpenAI findings →