// news · alignment · safety2026-08-19source: Lab publications

Two competing labs ran safety evaluations on each other's models

OpenAI and Anthropic completed a pilot cross-lab alignment exercise, each testing the other's models. Anthropic's assessment identified areas for improvement that broadly matched OpenAI's own research priorities.

OpenAI and Anthropic ran a pilot alignment evaluation exercise on each other's models — described as the first major cross-lab exercise of its kind — and published findings from it.

The most quoted result is the least surprising and the most encouraging: Anthropic's assessment flagged areas in OpenAI's models needing improvement that generally aligned with OpenAI's own active research priorities. Two rivals looking at the same system and independently agreeing on where it is weak is a small piece of evidence that these evaluations measure something real.

The precedent matters more than the findings. Safety assurance has run on self-assessment — a lab tests its own model and publishes a system card. Every incentive in that arrangement points one way. A competitor has the opposite incentive, which is precisely what makes the exercise worth something.

It is a pilot, run voluntarily, between two labs that already talk. It is not an audit regime: no obligation to participate, no consequence for declining, no independent party. The gap between "we agreed to test each other" and "someone may test us" is the whole distance to accountability.

It does land in a regulatory environment moving the same way. Authorities that can run their own evaluations rather than reading a provider's system card change who decides — and a voluntary version of that arriving first, from the labs, reads as an argument about how the mandatory version should look.

See our analysis →

OpenAI — Findings from a pilot Anthropic–OpenAI alignment evaluation exercise → · Anthropic Alignment Science — Agentic Misalignment in Summer 2026 →