Agent benchmarks are quietly bottlenecked on simulating the user
New work on interactive user-simulation toolkits and on evaluating human-agent systems under configurable human participation points at the same problem: most agent benchmarks either remove the human entirely or model them badly, and the gap between simulated and real users is where scores stop predicting deployment.
Removing the human is the standard simplification and it changes the task. Real deployments involve users who clarify, interrupt, change their minds and supply information only when asked. An agent optimised against a silent oracle is optimised against a counterfactual world.
Configurable human participation is the more useful framing, because participation is not binary. The interesting axis is how much the human supplies and when — and an agent that performs well at one level may be useless at another, which a single score cannot express.
This bears directly on the 2.6% figure. A pass rate measured without a realistic user is measuring autonomous completion, which is a harder task than most real deployments actually pose. The number is honest; what it is a number about needs stating carefully.
arXiv — VISTA: a versatile interactive user simulation toolkit for agent evaluation → · arXiv — HAS-Bench: evaluating LLM-based human-agent systems under configurable human participation → · arXiv — LLM reasoning is latent, not the chain of thought →