Why an Agent That Works 80% of the Time Finishes a Third of the Job
Eighty per cent sounds like a product you could ship. Most people hear it as "works, with the occasional miss" — annoying, survivable, the sort of thing a retry button covers.
Then it gets used on a flow with five steps in it.
The arithmetic nobody runs
ComponentBench
The authors make the point against their own results: a task requiring five critical interactions, each succeeding 80% of the time, has an end-to-end ceiling of roughly 33%. Their benchmark also found that 55.8% of failed traces end in a repeated-action loop, distributed evenly across every underlying cause — the loop is the symptom, not the fault.
Five steps at 80% is 0.8 to the fifth: about a third. Seven steps is 21%. Ten is 11%.
This is not a surprising piece of mathematics. It is just one that product decisions rarely get tested against, because reliability is usually quoted per action and experienced per task. The gap between those two numbers is where agent integrations quietly fail.
It is also the clearest argument for treating agentic experience design as a distinct practice. A flow that is pleasant for a person and eleven steps long is a flow an agent finishes one time in nine.
It also reframes what a step costs. An extra confirmation screen is not a small tax on an agent-driven flow — if the agent handles each screen at 80%, adding one takes a 33% success rate to 27%. The question stops being "is this step worth a click" and becomes "is this step worth a fifth of the completions."
Once is not the number that matters
Average success rates hide something worse: they tell you about one attempt, and real use is repeated.
τ-bench
A benchmark for agents working against real domain policies and tools, scored by comparing the final database state to an annotated goal state. GPT-4o succeeded on fewer than half of tasks — and its pass^8 score in the retail domain, meaning the same task completed correctly on all eight of eight attempts, fell below 25%.
Run the same task eight times. Under a quarter of the time does it work every time.
For a person deciding whether to keep delegating something, that is the figure that governs the decision. Nobody experiences your average. They experience a sequence, and one bad outcome in a sequence is what ends the habit — especially if the bad outcome was a duplicate order or a cancelled booking that should not have been.
A third of the wins are not wins
ST-WebAgentBench
222 tasks, each paired with explicit policy rules, scored across six dimensions: user consent, boundary and scope, strict execution, hierarchy adherence, robustness and security, and error handling. Measured as Completion Under Policy — crediting only completions that respected every applicable policy — three state-of-the-art agents averaged less than two-thirds of their nominal completion rate.
The agent finished. It also broke a rule on the way: acted outside the scope it was given, skipped a consent step, or carried on past something it should have stopped at.
Counting those as successes is how a system looks fine in testing and generates complaints in production. It is also a design brief. Each of those six dimensions is a question about your interface: where does this flow require explicit consent, and is that requirement visible to something that is not looking at the screen?
What compounding changes about agentic experience design
Shorten the critical path before improving any single step. Removing an interaction multiplies; making one interaction more reliable adds. If your checkout has nine steps, the fastest route to a working agent experience is seven steps, not a better date picker. The agent journey map is one way to count those steps across a whole task rather than a single screen.
Make operations idempotent and say so. Retrying is an agent's primary recovery strategy. If a retry can create a second order, every failed attempt becomes a support ticket rather than a second chance.
Give failures three distinguishable shapes. Wrong input, try again later, and never going to work need to look different in your responses. An agent that cannot tell them apart retries what will never succeed and abandons what would have worked a moment later.
Design one confirmable end state. The strongest thing you can offer is something the person can check afterwards — a reference number, a receipt, a link to the thing that changed. It is also the cheapest insurance against the pass^8 problem, because it converts an invisible failure into a visible one.
Where the numbers are soft
These figures come from benchmarks, and benchmarks measure a harness as much as a capability. WebArena's original 14.41% was later reached at 25.4% by a different team on the same tasks with better scaffolding. TheAgentCompany moved from 24% to 30% between paper versions. Anyone quoting a single number as the state of the art is quoting a date.
The compounding, though, is not a benchmark artifact. It is arithmetic, and it applies whatever the per-step rate turns out to be next year. Improving the rate moves the ceiling. Removing a step moves it faster.