HorizonMath arrives to measure AI progress toward genuine mathematical discovery — with automatic verification built in
As AI-generated proofs multiply, the field needs a way to measure them. HorizonMath is a new benchmark built to track AI progress toward mathematical discovery with automatic verification, so a claimed result can be machine-checked rather than taken on trust — the measurement apparatus a suddenly-productive field requires.
The benchmark exists because the old evaluation broke. Competition-math benchmarks measured whether a model could solve problems with known answers; that tells you nothing about discovery, where the answer is not known in advance. HorizonMath targets the harder question — can a system produce genuinely new, verifiable mathematics — and bakes automatic verification in so the measurement does not itself depend on trust.
Automatic verification is the load-bearing design choice. When results arrive faster than mathematicians can referee them, a benchmark that requires human review to score each entry cannot keep pace. Machine-checkable formalisation lets the benchmark validate a result's logic instantly, separating 'the proof is correct' from the slower human question of 'the result is interesting'.
The through-line to the cycle's headline is direct: Astra shipped Lean proofs precisely so they could be verified this way. The field is converging on formalisation as the common currency — the only way a flood of machine-generated mathematics can be trusted at the speed it is now being produced.
arXiv — HorizonMath: measuring AI progress toward mathematical discovery with automatic verification → · The Decoder — AI keeps cracking unsolved math problems, and mathematicians have mixed feelings →