// news · research-papers · alignment2026-08-11source: reporting

300,000 queries find frontier models disagree — and contradict their own specs

Researchers generated more than 300,000 queries probing value trade-offs across models from Anthropic, OpenAI, Google DeepMind and xAI. Each showed distinct prioritisation patterns, and the work surfaced thousands of cases of direct contradiction or interpretive ambiguity in published model specifications.

The finding about specifications is more actionable than the finding about disagreement. Labs publish documents stating how their models should weigh competing values; if a document contains thousands of internal contradictions and ambiguities, then "specification-compliant" is not a testable property, and a great deal of alignment evaluation quietly assumes it is.

That models differ is unsurprising and still worth having measured. Distinct, stable value orderings across labs mean the choice of provider is a policy choice embedded in a procurement decision — usually made by people who have never seen the spec.

Three hundred thousand queries is the kind of scale that only works automated, which brings the usual caveat: the generator's own priors shape the distribution of dilemmas. The contradictions found in the specs are robust to that concern, because a contradiction is a property of the document.

See our analysis →

Anthropic Alignment Science — Alignment Science Blog →