// blog · analysis · frontier-models2026-08-02source: quanta / the-decoder

The new frontier benchmark is a problem no one has solved

When top models saturate competition math, the yardstick has to change. In H2 2026 it changed to open conjectures — problems with no answer key. That reframing is the real capability story.

Cracking open conjectures has become the new frontier benchmark, displacing competition-math scores. The logic is simple: when models score near the ceiling on problems with known answers, the score stops discriminating. An open problem has no answer key, nothing to memorise, and an unambiguous success condition — it was open, now it is not.

Why this measures the real thing

Research mathematics demands exactly the long-horizon, verifiable reasoning that also underwrites high-value agentic and coding work. A model that moves an open problem is demonstrating the underlying capability, not just a party trick. That is why the labs are racing on discovery rather than chat quality — it is the hardest honest test of the thing they are all trying to build.

But the field is right to stay measured. Terence Tao compared the moment to the early-20th-century foundational crisis — disorienting, forcing a rethink of authorship and proof, but plausibly generative rather than destructive. A handful of dramatic results, several unrefereed, is a signal, not a settled capability.

The reframing is the story

Whether or not every headline result survives review, the benchmark has moved. 'Can it move an open problem' has replaced 'can it ace a test,' and that shift in what we ask of a frontier model is itself the most durable thing that happened this quarter.

Quanta Magazine — The AI revolution in math has arrived → · The Decoder — AI keeps cracking unsolved math problems →