The new frontier benchmark is a problem no one has solved
When top models saturate competition math, the yardstick has to change. In H2 2026 it changed to open conjectures — problems with no answer key. That reframing is the real capability story.
Cracking open conjectures has become the new frontier benchmark, displacing competition-math scores. The logic is simple: when models score near the ceiling on problems with known answers, the score stops discriminating. An open problem has no answer key, nothing to memorise, and an unambiguous success condition — it was open, now it is not.
Why this measures the real thing
Research mathematics demands exactly the long-horizon, verifiable reasoning that also underwrites high-value agentic and coding work. A model that moves an open problem is demonstrating the underlying capability, not just a party trick. That is why the labs are racing on discovery rather than chat quality — it is the hardest honest test of the thing they are all trying to build.
But the field is right to stay measured. Terence Tao compared the moment to the early-20th-century foundational crisis — disorienting, forcing a rethink of authorship and proof, but plausibly generative rather than destructive. A handful of dramatic results, several unrefereed, is a signal, not a settled capability.
The reframing is the story
Whether or not every headline result survives review, the benchmark has moved. 'Can it move an open problem' has replaced 'can it ace a test,' and that shift in what we ask of a frontier model is itself the most durable thing that happened this quarter.
Quanta Magazine — The AI revolution in math has arrived → · The Decoder — AI keeps cracking unsolved math problems →