Agent benchmarks are being quoted by the vendors who score on them
Browser Use reports 89.1% on WebVoyager. TinyFish reports 91.1%. OpenAI's Computer-Using Agent reports 87% on the same set and 58.1% on WebArena. Almost every headline agent number in circulation is self-reported, and the 30-point gap between the two benchmarks is the more useful fact.
The agent numbers being quoted this month are striking, and almost all of them come from the vendor being measured. Browser Use reports 89.1% on WebVoyager. TinyFish reports 91.1% across a unified platform. OpenAI reports 87% on WebVoyager for its Computer-Using Agent — and, in the same disclosure, 58.1% on WebArena.
That last pairing is the most informative line in the set. The same system scores 87% on one web-agent benchmark and 58.1% on another. The benchmarks are not measuring the same thing: WebVoyager runs live sites with a model-judged success criterion; WebArena runs reproducible self-hosted environments with programmatic checks. The harder, more verifiable one produces the lower number.
OSWorld sits in a third position again — 369 real computer tasks across browsing, file management, spreadsheets and cross-application workflows — and is the set most directly relevant to anyone evaluating desktop automation rather than web navigation.
The adoption figures are moving regardless. McKinsey reports 62% of organisations experimenting with or deploying agents; Gartner projects that 40% of enterprise applications will include task-specific agents by the end of 2026, against fewer than 5% in 2025. Buying decisions at that scale are being made against self-reported scores on benchmarks that disagree with each other by thirty points.
Firecrawl — 11 Best AI Browser Agents in 2026 → · TinyFish — Best AI Browser Agents in 2026: Compared by Use Case → · Automation Anywhere — AI Agent Benchmarks: The 2026 Enterprise Evaluation Guide →