Agent Experiences

Designing the Agent Handoff

By Alan WeibelPublished 9 min read

Almost everything written about agent experience assumes the agent finishes. The interesting design problem is the other case, and it is the one nobody has a pattern for: the agent is halfway through, it is wrong or uncertain or blocked, and a person has to be brought back in.

That moment has a cost, a success rate, and a set of decisions behind it. There is research on all three, and it lands squarely in agentic experience design: the handoff is an interface, and almost nobody has designed theirs.

A little help goes a very long way

Magentic-UI

A human-in-the-loop agent system built around six interaction mechanisms, including co-planning, action approvals and answer verification. On 162 GAIA validation tasks, adding a knowledgeable simulated user raised completion from 30.3% to 51.9% — a 71% relative improvement. The agent asked for help in only 10% of tasks, and when it did, an average of 1.1 times.

Source: Mozannar et al., Microsoft Research, 2025

Ten per cent of tasks. One question each. Half as many failures.

That ratio is the entire economic argument for designing a handoff rather than chasing autonomy. The expensive thing is not human involvement; it is human involvement at the wrong moments, or none at the moment that mattered.

The same study is honest about the cost of getting it wrong. Twelve participants used the system for an hour each. It scored 74.58 on the System Usability Scale and 91.7% did not find it unnecessarily complex — but only 41.7% agreed they would like to use it frequently. The friction people complained about was latency, verbosity, and errors, and they were "uncertain about system capabilities," building a mental model by trial and error.

The authors also name the open problem plainly: they have no ground-truth signal for when to interrupt. That is the unsolved part, and it is a design problem before it is a modelling one.

Knowing when to stop is a distinct failure

Computer-use agents for blind users

A three-week diary study with eight blind participants issuing 1,258 commands across 12 desktop applications, with full traces captured. GPT-5 achieved the highest success rate at 52.5%. Trace analysis sorted the failures into four classes: grounding, planning, constraint-tracking, and termination — the last being the agent not knowing when it was finished.

Source: Kodandaram et al., EMNLP 2026

Termination failure is the one worth staring at, because it is the one your interface causes.

An agent that cannot tell whether it is done will either stop early and report success it did not achieve, or keep going past the point it should have stopped. Both are products of an application that never says, unambiguously, that the thing happened. A spinner that disappears is not a confirmation. A page that navigates is not a receipt.

The oversight literature disagrees with itself, usefully

Here the research stops converging, and the disagreement is more useful than either side alone.

Cognitive forcing functions

Harvard researchers tested three cognitive forcing designs against simple explainable-AI approaches with 199 participants. Forcing people to engage before seeing the AI's answer significantly reduced overreliance — but participants gave the least favourable subjective ratings to exactly the designs that helped most. The benefit was also larger for people high in Need for Cognition, raising a fairness concern.

Source: Buçinca et al., CSCW 2021

Explanations can reduce overreliance

Stanford and UW researchers ran five studies with 731 participants and argued overreliance is not a fixed property of human cognition but a cost-benefit decision. Manipulating task difficulty, explanation difficulty and monetary benefit each changed how much people over-relied — people engage with an explanation when engaging is worth it.

Source: Vasconcelos et al., CSCW 2023

One says friction works but people dislike it. The other says the problem is not human laziness, it is that checking costs too much — so lower the cost.

Both are actionable and they point at different work. From the first: expect the design that improves outcomes to test badly in usability research, and do not optimise it away on the strength of a satisfaction score. From the second: make the diff cheap to read, the trace scannable, the claim easy to spot-check. Most products have far more room on the second than on the first.

Confidence goes up as correctness goes down

Do users write more insecure code with AI assistants?

Stanford researchers found participants with access to an AI assistant "wrote significantly less secure code than those without access" — and were more likely to believe their code was secure. Participants who trusted the assistant less, and engaged more with their prompts, produced fewer vulnerabilities.

Source: Perry et al., ACM CCS 2023

This is the cleanest demonstration of automation bias with a hard outcome measure rather than a self-report, and the inversion is the design target. People were more confident and less correct at the same time.

Anything in your interface that conveys an agent's actual uncertainty is doing safety work. Anything that smooths it over — a confident summary, a tick, a "done" with no detail — is doing the opposite.

Six agentic experience design questions for the handoff

Magentic-UI's six mechanisms translate almost directly into things to design:

  1. Co-planning. Can a person see and change the plan before it runs, or only the result after?
  2. Action approval. Which actions stop and ask? Irreversible ones, ones that spend money, ones that touch someone else.
  3. Answer verification. What can the person look at to check the claim, without redoing the work?
  4. Interruption. Can a person step in mid-task, or only cancel?
  5. Memory. Does a correction stick, or will the next run make the same mistake?
  6. The return. When the agent gives up, what does the person receive? A dead end, or a state they can continue from?

The sixth is the one most often missing, and it is the one that decides whether somebody tries again. Those last two — recovering and handing back — are the closing steps of the agent journey map for software, and they are where most teams discover they have designed a path with no exit.

What this does not tell you

The strongest numbers here come from a simulated user, not a real one — a knowledgeable model standing in for a person, which flatters the result. The qualitative work behind it is twelve participants, which is a reasonable size for finding problems and much too small for measuring their frequency.

What survives that is the shape rather than the magnitude: help is worth far more than its cost, the hard part is knowing when to ask, and a product that communicates its own uncertainty does better by the person using it than one that performs confidence.