// news · tools · agents2026-08-21source: Benchmark documentation

OSWorld is the benchmark that matches what buyers actually automate

369 real computer tasks across browsing, file management, spreadsheets and cross-application workflows. For anyone evaluating desktop automation rather than web navigation, it is the most directly relevant public number — and the least quoted in marketing.

OSWorld evaluates computer-use agents across 369 real computer tasks spanning web browsing, file management, spreadsheet manipulation and cross-application workflows. For enterprises evaluating agents for desktop automation, it is the most directly relevant public benchmark available.

It is also conspicuously less quoted than WebVoyager. The reason is not mysterious: cross-application workflows are where agents break. Moving data from a spreadsheet into a web form into a file with a particular name is exactly the kind of task where a small error compounds silently, and the scores reflect that.

That makes it the useful one. A buyer is not automating the median browsing task, they are automating a workflow that touches four applications and has a correct answer someone downstream depends on. The benchmark that produces uncomfortable numbers on that shape of work is more informative than the one that produces flattering numbers on navigation.

The practical guidance for evaluation is to ask vendors for OSWorld alongside whatever they lead with, and to treat a refusal as data. Where the same system scores 87% on one set and 58.1% on another, which set the vendor chooses to publish is itself informative.

See our analysis →

Automation Anywhere — AI Agent Benchmarks: The 2026 Enterprise Evaluation Guide → · Bright Data — 10 Best Agentic Browsers for AI Automation in 2026 →