// news · research-papers · alignment2026-08-13source: arXiv

A Benchmark Health Index proposes scoring the evaluations themselves

The Benchmark Health Index sets out a systematic framework for benchmarking the benchmarks of large language models — treating an evaluation as an instrument with measurable properties like saturation, contamination exposure, discriminative power and reproducibility, rather than as ground truth.

The field has spent four years arguing about model scores and almost no time asking whether the rulers are any good. A health index for benchmarks is overdue infrastructure, and its value is less in any particular ranking than in making "this benchmark is worn out" a claim you can support with numbers.

The properties it makes measurable are the ones practitioners already gossip about: saturation, where the top of the range compresses until differences stop being meaningful; contamination exposure, where a benchmark's age and popularity make training-set leakage likely; discriminative power, whether the test separates models that genuinely differ; reproducibility, whether the same evaluation run twice gives the same answer.

The systemic problem it addresses is that benchmarks are load-bearing far outside research. They appear in launch materials, procurement decisions, safety cases and — increasingly — regulatory documentation. A number that cannot be shown to be reliable is doing a lot of work in places where reliability is the entire point.

Pair it with benchmarks that move when you rephrase the question and the message is consistent: the evaluation layer is now the least rigorous part of a field that measures everything.

See our analysis →

arXiv — Benchmark Health Index: A Systematic Framework for Benchmarking the Benchmarks of LLMs → · arXiv — PaperBench: Evaluating AI's Ability to Replicate AI Research →