// news · research · agents2026-08-17source: arXiv

Evaluating agents on nine benchmarks costs about $40,000

The Holistic Agent Leaderboard's own accounting puts a single full evaluation across nine benchmarks at roughly $40,000. Evaluation has become a capital expense, and that changes who gets to check anyone's work.

Running the Holistic Agent Leaderboard across its nine benchmarks costs approximately $40,000 by the authors' own accounting.

That figure deserves to be quoted more often than it is, because it settles an argument about independent verification. At $40,000 a pass, evaluating agents is not something a graduate student does on a whim, and it is not something a journalist does at all. The parties who can afford to check a claim are the parties with an interest in the claim.

It also compounds badly with velocity. Nine benchmarks across coding, web navigation, general assistance and customer service is a reasonable spread, and every new agent scaffold requires the whole thing again. A field that ships weekly cannot afford to verify weekly at this price.

Which is why the cost-reduction work matters more than it sounds. A companion result finds that small task subsets preserve rank order, and if that holds it turns a $40,000 question into a manageable one.

The uncomfortable version of this: the reason agent claims are hard to check is not that the methods are secret. It is that checking is expensive, and expense is a more durable barrier than secrecy.

See our analysis →

arXiv — Efficient Benchmarking of AI Agents → · arXiv — Benchmark Test-Time Scaling of General LLM Agents →