Shipped is not the same as trusted
Gartner forecasts 40% of enterprise applications will include a task-specific agent by the end of 2026. McKinsey reports 62% of organisations already experimenting. Neither number tells you whether anyone lets an agent finish a task unsupervised.
Gartner projects 40% of enterprise applications will include task-specific agents by the end of 2026, up from under 5% in 2025, while McKinsey reports 62% of organisations already experimenting or deploying. Both figures are probably defensible. Neither measures the thing that determines whether this wave succeeds.
Three different questions
“Includes an agent” is a shipping fact about a vendor’s product roadmap. A feature flag satisfies it. “Experimenting or deploying” is a survey response, and the two halves of that phrase are separated by roughly the entire distance between a pilot and a production system. Neither is adoption, and adoption is not the same as reliance.
The number that matters is the fourth one nobody publishes: what fraction of tasks an agent starts are completed without a human intervening. That is the figure that determines whether the agent removes work or relocates it, and it is the figure that decides whether the 2027 stories are about transformation or disillusionment.
Why the gap manufactures disappointment
The mechanism is predictable and already visible in prior enterprise-software cycles. Vendors ship the capability, analysts count the shipping, buyers read the count as evidence of maturity, deploy against that expectation, and discover that the supervision burden is where the cost lives. The technology did not fail. The measurement described something other than what the buyer thought they were buying.
The benchmark that would help, and why it is ignored
There is a public evaluation that maps unusually well onto enterprise reality. OSWorld covers 369 real computer tasks across browsing, file management, spreadsheets and cross-application workflows — which is a fair description of what a back-office automation project actually consists of. It is also among the least quoted numbers in agent marketing, for the obvious reason that scores on it are low.
That is precisely what makes it useful. A benchmark everyone scores 90% on cannot inform a purchase. A hard benchmark aligned to the real workload can, and the discomfort it causes is the information.
What to ask before deploying
What is the unassisted completion rate on our workflows, measured by us. What happens on failure — does the agent stop, or does it proceed confidently and wrongly. Who reviews the output, and is that person cheaper than the person the agent replaced. And what the fallback is when the model behind the feature is updated without notice.
Those questions are answerable, and answering them takes weeks rather than a quarter. The organisations that ask them will end up in the 40% too. They will just know what they bought.
Automation Anywhere — AI Agent Benchmarks: The 2026 Enterprise Evaluation Guide → · Bright Data — 10 Best Agentic Browsers for AI Automation in 2026 →