The Same Agent, the Same Page, 34 Points Apart
ComponentBench
Researchers at Duke and Amazon's AGI SF Lab tested seven models against 97 canonical UI component types across 2,910 verified tasks, built on Ant Design, MUI and Mantine. Holding the model and the task constant and changing only what the agent was allowed to observe, GPT-5 mini scored 83.1% working from an accessibility tree and 48.9% working from coordinates on a screenshot.
Most advice about building for AI agents starts with the model: which one, how good is it, what will the next one fix. That framing survives right up until someone runs the experiment where the model does not change.
It is also the framing that leaves designers with nothing to do, which is why this result belongs in agentic experience design rather than in a machine learning roadmap.
Change only the representation the agent is given — the accessibility tree, or pixels and coordinates — and the same model on the same tasks moves by more than thirty points. The intelligence was constant. What varied was whether the page explained itself.
What the gap is made of
The ComponentBench authors break 8,864 failed traces into causes, and the top of the list is not reasoning. It is continuous calibration error at 20.2%: the agent knew what it wanted and could not hit it. Then transient state loss at 19.9%, a missing commit or confirmation step at 11.6%, and target acquisition — clicking the wrong instance of the right thing — at 11.2%.
Those are not failures of judgement. They are failures of aim, and aim is a property of the interface.
The component-level results say the same thing more bluntly. Averaged across six models:
| Component family | Pass rate |
|---|---|
| Command and navigation | 91.6% |
| Overlays and transient UI | 89.2% |
| Discrete choice | 86.3% |
| Disclosure and progressive | 72.1% |
| Date and time | 71.2% |
| Advanced editors | 61.9% |
| Continuous precision | 59.4% |
| Drag, drop and workspace | 47.7% |
A drag-to-reorder is close to a coin flip. A menu is close to free.
The inversion worth sitting with
Nine component types in that study take a human two steps or fewer and score under 60% for agents. A range slider: 1.9 human steps, 39.9% for agents. A window splitter: 1.3 steps, 38.3%. Resizable columns: 1.7 steps, 24.4%.
And a native select element — 1.4 human steps — scores 53.1%.
No component in the study shows the reverse pattern. There is nothing that is hard for people and easy for agents. The difficulty curve only bends one way, which means intuitions built from watching users do not transfer. The control you reach for because it feels effortless may be the one that stops an agent dead.
This is the part of designing for agents that separates it from UX rather than extending it. A usability finding tells you a control was confusing and a person recovered anyway. Here there is no recovery, because there is nobody to improvise.
WorkArena
ServiceNow Research and Mila tested agents on 33 tasks drawn from real ServiceNow instances, covering 19,912 task instances. On the six list-filtering tasks, GPT-4o, GPT-4o-V, GPT-3.5 and Llama3-70B each scored 0.0%. The authors attribute it to a non-standard HTML widget and note the tasks are "fairly simple for a human to complete."
One widget, every model, zero
The WorkArena result is the one to show a stakeholder who wants a number.
Four models. Six tasks. Zero per cent — not a low score, an empty column. The same agents that completed 77.8% of service catalogue tasks and 80% of knowledge base tasks could not filter a list, because the filter was built as a custom widget rather than from standard elements.
A human barely notices the difference. The control looks like a filter, behaves like a filter, and has a little funnel icon on it. To anything reading the markup it is a div with handlers attached, and there is nothing to operate.
The same paper measures something else worth knowing: the flat HTML of those pages runs between 40,000 and 500,000 tokens after scripts, styles and empty elements are stripped. The authors use accessibility trees instead, calling the raw HTML prohibitively large. Page weight is not only a performance problem anymore. It is a comprehension budget.
The half agentic experience design can fix
SeeAct
Ohio State researchers found GPT-4V completed 51.1% of tasks on live websites when its textual plans were manually grounded into actions by a human — that is, when a person did the work of finding the right element on the page. Their conclusion was that "grounding still remains a major challenge."
Split an agent's job in two. There is deciding what to do, and there is finding the thing to do it to. SeeAct isolates the second by having humans do it, and performance roughly doubles.
That division matters because the two halves belong to different people. The first half is the model vendor's problem and it improves without you. The second half is your markup, and it does not improve unless you improve it.
What to change in your interface
Nothing here is exotic. It is close to the accessibility advice that has been on this site for a year, which is the point — the two audiences need the same things for different reasons.
Prefer standard elements over custom ones. A native select beat a custom
widget in both studies. Where a custom control is genuinely necessary, give it
the roles, states and labels a standard one would have had.
Count the required interactions in your critical flows. Each one is a place to fail, and they multiply rather than add. More on that in the next piece.
Treat anything continuous as hostile. Sliders, splitters, drag handles and colour pickers are the bottom of every table above. Where one is the only way to set a value, provide a second way that takes a number.
Look at your page the way an agent does. Not the rendered view — the accessibility tree, which is what the better agents read. Your browser's dev tools will show it. If a control is missing from that tree, it does not exist.
For the interaction-level decisions underneath these — which control to reach for and when — our AXD interaction patterns cover the reusable ones.
The caveat these numbers deserve
Benchmark figures age badly and the field has a habit of quoting them long after they are wrong. A 2025 paper from Ohio State and Berkeley, pointedly titled An Illusion of Progress?, argues that previously reported results overstate what agents can do, largely because of how the benchmarks are built.
Take the absolute percentages here as of their publication dates and no later. The durable finding is not any single number. It is the direction every one of these studies points: the representation you expose decides the outcome, and it is the part of the problem that is yours.