// blog · analysis · research-papers2026-08-11source: ICLR 2026 / reporting

Benchmarking the scientist

Most agent leaderboards cannot tell you whether the winner had a better model or just a more expensive one. Fixing that is unglamorous and it is the whole job.

AstaBench covers 2,400+ problems across the scientific discovery process, holding tooling and cost constant.

The confound problem

An agent scoring higher may have a stronger base model, a larger inference budget, richer tools, or a harness quietly tuned to the benchmark. Published results rarely separate these, which means a leaderboard position is often a procurement decision disguised as a capability claim.

Standardising interfaces and controlling for cost is tedious work that generates no headlines and makes every subsequent number mean something. That trade is worth taking.

Lifecycle over task

Real research is search, then hypothesis, then method, then analysis, then writing. An agent that is excellent at literature review and useless at experimental design produces nothing at all — but scores well on a benchmark that tests stages in isolation. Covering the chain punishes exactly the profile that isolated benchmarks reward.

The limit nobody escapes

To score an answer you must already know it. AstaBench measures whether an agent can retrace paths humans have walked. That is necessary and it is not discovery, and the gap between the two is where every interesting claim about AI-for-science lives.

There is a second-order finding worth pairing with it. 300,000 queries probing value trade-offs surfaced thousands of contradictions inside labs' own published specifications. Benchmarks assume a stable target. Some of the targets are not stable documents.

ICLR 2026 Proceedings — AstaBench → · Anthropic Alignment Science — Alignment Science Blog →