// blog · analysis · research-papers2026-08-17source: arXiv

The price of knowing

Checking an agent claim across nine benchmarks costs about forty thousand dollars. Expense is a more durable barrier to verification than secrecy, and nobody put it there on purpose.

A full pass of the Holistic Agent Leaderboard across nine benchmarks costs approximately $40,000, by the authors' own accounting.

Who can afford to check

Not a graduate student. Not a journalist. Not a customer evaluating a purchase. At that price the parties able to verify an agent claim are overwhelmingly the parties with an interest in the claim, and that is a structural problem rather than a bad-actor problem.

The reason agent claims are hard to check is not that the methods are secret. It is that checking is expensive.

And it compounds with velocity

Nine benchmarks spanning coding, web navigation, general assistance and customer service is a reasonable spread. Every new scaffold requires the whole thing again — and agent results are unusually sensitive to scaffolding rather than to the model. A field that ships weekly cannot verify weekly at this price.

Which is why the cheap version matters

Across eight benchmarks and 33 scaffolds, rank order stayed stable on small task subsets. The protocol needs no optimisation: evaluate a new agent only on tasks with intermediate historical pass rates.

The selection rule is the elegant part. Tasks everything passes carry no information; tasks everything fails carry none either. The middle discriminates, and the middle can be read off history rather than searched for.

What it does not give you

Rank order, not absolute scores. And the question a buyer actually asks is absolute — can this do my task, not is it better than that one. A cheap ranking is a real contribution and it is not a cheap evaluation.

Still, it is the field trying to make its own claims checkable by someone other than the claimant. That instinct is also visible in the index of what is actually deployed, and it is the healthiest thing happening in agent research right now.

arXiv — Efficient Benchmarking of AI Agents → · arXiv — Benchmark Test-Time Scaling of General LLM Agents → · arXiv — Benchmarking Agentic Code Reasoning at the Repository Level →