// blog · analysis · research-papers2026-08-13source: arXiv

Benchmarking the benchmarks

Rewrite an AIME problem thirteen ways without changing the mathematics and the scores move. A field that measures everything has left its rulers unmeasured.

The Robust Reasoning Benchmark applies thirteen deterministic perturbations to AIME 2024 and 2025. Separately, a Benchmark Health Index proposes scoring evaluations as instruments.

The unstated assumption

Every benchmark result carries one: that the score measures the capability rather than the phrasing. It is almost never tested, because testing it is unglamorous and the result is unflattering.

RRB tests it directly, and the word doing the work is deterministic. The thirteen perturbations are reproducible transformations, not paraphrases sampled from another model, so the variance they produce is attributable to the model under test and not to the rewriter. Rename variables, reorder independent clauses, change surface presentation — none of it should matter to a system that solved the problem.

Instruments have properties

The health-index framing is the more portable contribution. It makes measurable the things practitioners already say informally: saturation, where the top of the range compresses until differences stop meaning anything; contamination exposure, where age and popularity make leakage likely; discriminative power; reproducibility.

"This benchmark is worn out" becomes a claim you can support with numbers.

Why it matters outside research

Because benchmarks are load-bearing far beyond papers. They appear in launch materials, procurement decisions, safety cases and increasingly in regulatory documentation. A composite score that says parity while a sub-score says a third less capability is not a research curiosity when someone is choosing an agent platform on it.

The practical takeaway for reading any model card: a headline number measures performance on a specific set of sentences. How much survives rephrasing is a different number, and almost nobody reports it.

arXiv — Robust Reasoning Benchmark → · arXiv — Benchmark Health Index: A Systematic Framework for Benchmarking the Benchmarks of LLMs → · arXiv — ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning →