Frontier benchmark leadership splits: OpenAI tops science reasoning, Anthropic leads real-world coding
There is no single frontier leader anymore. On GPQA Diamond — graduate-level science reasoning — GPT-5.4-Pro leads at 94.4%; on SWE-Bench Verified — real-world software engineering — Claude Opus leads at 87.6%. The crown has split by domain, and buyers now pick a model per task rather than a single best one.
The split is what a maturing frontier looks like. When one lab tops science reasoning and another tops software engineering, 'best model' stops being a meaningful single ranking and becomes a per-domain question. The benchmarks that matter diverge, and so do the leaders on them.
For buyers this is a procurement shift, not a curiosity. A team doing scientific analysis and a team shipping code now have different optimal models, which pushes organisations toward multi-model routing — a gateway that sends each task to whichever model leads its domain, rather than standardising on one vendor. The abstraction layer, not the model, becomes the durable choice.
The strategic consequence for the labs is that domain leadership is defensible in a way overall leadership is not. Owning software engineering, or science reasoning, or long-context outright is a clearer moat than a fleeting top-of-the-average position — which is why the labs increasingly optimise for specific high-value benchmarks rather than a single composite score.
David Veksler — Frontier AI labs list: companies, models and strategy (2026) → · LLM-Stats — AI updates today (August 2026) →