AstaBench scores agents across 2,400+ problems spanning the scientific discovery process
The benchmark targets the failures that make agent evaluation unreliable: non-reproducible tooling, confounds from model cost and tool access, non-standard interfaces, and tasks that do not resemble real work.
The confound problem is the reason most agent leaderboards are hard to read. An agent scoring higher may have a better model, more expensive inference, richer tools, or a harness tuned to the benchmark — and published results rarely separate those. Holding tooling and cost constant is unglamorous and it is what makes a number mean something.
Covering the whole lifecycle rather than isolated tasks is the other useful choice. Real research is literature search, then hypothesis, then method, then analysis, then writing, and an agent that is strong at one stage and useless at the next produces nothing. Stage-wise benchmarks reward exactly that profile.
The standing limitation of any benchmark aimed at discovery is that the answers must already be known to be scored. AstaBench measures whether an agent can retrace paths humans have walked — necessary, not sufficient, and the gap between the two is where the interesting claims live.