// blog · analysis · agents2026-08-21source: Vendor benchmark disclosures and agent security literature

The agent numbers are marketing until someone else runs them

Browser Use says 89.1% on WebVoyager. TinyFish says 91.1%. Both are self-reported. The interesting number is not either of those — it is the thirty-point drop when the same class of system is measured on a harder set.

Almost every headline agent capability figure in circulation was produced by the vendor whose product it flatters. Browser Use reports 89.1% on WebVoyager; TinyFish reports 91.1%; OpenAI’s Computer-Using Agent reports 87% on the same set and 58.1% on WebArena. That last pair is the only comparison in the list that is internally controlled, and it is the one worth staring at.

Self-reporting is not the same as lying

It is worth being precise about the failure mode, because the accusation of dishonesty is both unfair and less useful than the truth. Vendors running their own benchmarks are usually running them correctly. What they control is everything around the number: which scaffold, how many attempts, what counts as success on a partially completed task, which subset was reported, and how many configurations were tried before this one was published. None of that requires bad faith. All of it moves the number.

The reproducibility problem is that no one else can rerun it. Agent evaluations depend on a live web that changes underneath the test, on scaffolds that are frequently unreleased, and on grading criteria that are often a sentence long in a blog post.

The thirty points

87% on WebVoyager and 58.1% on WebArena, same system, same evaluator. That spread is not noise and it is not a vendor artefact — it is a statement about how much benchmark difficulty varies within what everyone calls “web agents.” Any procurement decision made on a single agent benchmark is implicitly betting that the buyer’s workload resembles that benchmark’s difficulty. Thirty points is the size of the bet.

Where the field is going instead

The more promising development this month is not a benchmark at all. A framework paper argues that agent security cannot be addressed at the model layer because the failure modes are distributed across architecture, deployment and operation. That is the same insight applied to reliability: the thing that determines whether an agent works in production is mostly not the model, and a benchmark that varies only the model is measuring the wrong variable.

What to ask a vendor

Which benchmark, run by whom, with what scaffold, at how many attempts, and what the same system scores on a second benchmark of different difficulty. The fifth question is the one that separates vendors who have measured their system from vendors who have marketed it.

Firecrawl — 11 Best AI Browser Agents in 2026 → · TinyFish — Best AI Browser Agents in 2026: Compared by Use Case → · Automation Anywhere — AI Agent Benchmarks: The 2026 Enterprise Evaluation Guide → · arXiv — Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability →