A Benchmark Health Index, and a detective game whose usefulness expires this year
One paper proposes a systematic framework for benchmarking the benchmarks of LLMs. Another builds a naturalistic reasoning benchmark from a tabletop detective game, watches models climb from the lower quartile of human performance to the top 14 percent in nine months, and states plainly that its utility is nearly over.
The Watson and Holmes result is the more vivid of the two. A benchmark that took nine months to go from below-average-human to top-14-percent, with saturation expected by the end of the year, has a shorter useful life than the review cycle of the paper describing it.
That is what makes the Benchmark Health Index necessary rather than academic. If individual benchmarks expire this fast, the field needs a way to assess whether a given measurement is still measuring anything — contamination, saturation, construct validity — as a standing property rather than a one-time check.
It also reframes the sandbox story. A model that reached GitHub mid-evaluation is a contamination event, and contamination is precisely what a health index is supposed to catch before the score is published rather than after.
arXiv — Benchmark Health Index: a systematic framework for benchmarking the benchmarks of LLMs → · arXiv — Watson & Holmes: a naturalistic benchmark for comparing human and LLM reasoning → · arXiv — ARC-AGI-2: a new challenge for frontier AI reasoning systems →