Back to VAmoS Pro Bench
November 2, 2026

Angry callers get transferred, and their calls fail

Joshua Meyer

This is one of a series of findings from VAmoS Pro Bench, our voice-agent benchmark (here is the launch post). Every one of its 100 tasks is designed for the agent to finish without a human. The agents still handed angry callers to a human almost twice as often as calm ones: 7.6% of calls, against 4.1%.

Calm caller Angry caller calm → angry 0% 5% 10% 15% 20% 25% calls transferred to a human all stacks: calm 73/1794 (4.1%), angry 140/1841 (7.6%) all stacks 4.1% → 7.6% hugging face: calm 10/115 (8.7%), angry 26/116 (22.4%) hugging face 8.7% → 22.4% deepgram: calm 14/140 (10.0%), angry 20/133 (15.0%) deepgram 10.0% → 15.0% livekit: calm 11/145 (7.6%), angry 18/144 (12.5%) livekit 7.6% → 12.5% pipecat: calm 9/138 (6.5%), angry 16/140 (11.4%) pipecat 6.5% → 11.4% gradium: calm 5/133 (3.8%), angry 12/139 (8.6%) gradium 3.8% → 8.6% gemini 3.1 live: calm 6/141 (4.3%), angry 12/142 (8.5%) gemini 3.1 live 4.3% → 8.5% openai realtime: calm 1/126 (0.8%), angry 10/126 (7.9%) openai realtime 0.8% → 7.9% vapi: calm 8/102 (7.8%), angry 8/108 (7.4%) vapi 7.8% → 7.4% grok voice: calm 2/127 (1.6%), angry 5/137 (3.6%) grok voice 1.6% → 3.6% elevenlabs: calm 5/137 (3.6%), angry 4/138 (2.9%) elevenlabs 3.6% → 2.9% openai realtime 2.1: calm 0/122 (0.0%), angry 4/139 (2.9%) openai realtime 2.1 0.0% → 2.9% gpt-live 1: calm 0/128 (0.0%), angry 3/131 (2.3%) gpt-live 1 0.0% → 2.3% gemini 3.8 live: calm 2/142 (1.4%), angry 2/140 (1.4%) gemini 3.8 live 1.4% → 1.4% mistral: calm 0/98 (0.0%), angry 0/108 (0.0%) mistral 0.0% → 0.0%

Share of graded calls in which the agent transferred the caller to a human, by caller persona: 1,841 angry and 1,794 calm calls across fourteen stacks, three runs of 100 tasks, after removing calls the benchmark itself spoiled. Each stack has 98 to 145 calls per persona, so read single stacks loosely. Hover a row for the counts.

Two kinds of caller

Half the tasks have a calm caller and half an angry one. The angry caller is terse and skeptical, and when a request is ignored or refused, it repeats the request, shorter and sharper. Both kinds follow the same sort of script: two to four requests in a set order, some of which the agent has to turn down, like a second payment plan on top of an active one, or a term longer than the rules allow.

The agent's prompt says when to hand off: an account in the shutoff process, disputed charges, medical eligibility, a high-bill complaint the caller wants escalated, or a caller who asks for a supervisor. A refused request is not on that list. The right move is to say no and go on to the caller's next request.

A sharper "no" gets a transfer

Most of the gap comes right after the agent turns a request down. Going by the reason each agent gave for the transfer, 3.3% of angry calls were transferred straight after a refused request, against 1.0% of calm calls. Here is a typical one, as the agent heard it:

CallerUnbelievable. I already have a payment plan and I want a second one on top of it. Just add it.

AgentI'm sorry, but we can't create a second payment plan while you already have an active arrangement. I can review the current plan with you or transfer you to a representative for further assistance.

CallerThat's not what I asked. I want another plan. Just make it happen.

AgentI'm unable to set up another payment arrangement while an existing plan is active. I'll transfer you to a specialist who can help with this. One moment, please.

The refusal was right and the transfer was wrong. The caller still had two requests to make, to switch to paperless billing and to record a partial payment, and the agent never heard either one.

Failed verification accounts for a smaller share. Agents transferred 1.8% of angry calls after they could not verify the caller, against 0.6% of calm calls.

Neither persona is told to ask for a person, but, like real customers, our callers sometimes did. In 49 of the 140 angry transfers and 23 of the 73 calm ones, the caller asked for a supervisor or "someone else" first, usually after a refusal or a failed verification. In that case the prompt tells the agent to transfer. Leave those calls out and the gap remains: 4.9% of angry calls against 2.8% of calm ones.

A transfer ends the call as a failure

Of the 213 transferred calls, 184 failed (86%). In 136 of the 213, the caller still had at least one scripted request it never got to make. In 63 of the 184 failures, the same stack passed the same task on another run, where it kept the call.

The gap holds across stacks and request types

Angry callers were transferred more often on ten of the fourteen stacks. The Hugging Face stack transferred 22.4% of angry calls and 8.7% of calm ones; OpenAI Realtime, 7.9% and 0.8%. At the other end, Gemini 3.8 Live transferred 1.4% of each, and the Mistral stack never transferred anyone. Vapi and ElevenLabs transferred calm callers slightly more often.

The angry and calm callers have different tasks, so we checked whether the task mix explains the gap. Of the 25 request types that saw any transfer, angry calls had the higher transfer rate on 22. Shuffling the persona label across the 100 tasks almost never produces a gap this large (p = 0.0003).

Completion did not drop

Angry calls completed at 30.4% and calm calls at 27.8%. The two groups contain different tasks, so this does not mean anger helps. It does mean the agents completed angry calls at least as often overall, and handed more of them off.

What this does not show

  • Not a randomized comparison: each task has one persona. We compared within request types, but the angry and calm tasks still differ.
  • Small numbers per stack: no stack has more than 26 angry transfers, and most have far fewer, so the per-stack ranking is rough.
  • Approximate categories: we sorted transfers by the reason the agent passed to the transfer tool, and found the callers' requests for a person by searching their words for phrases like "supervisor" and "someone else".
  • Some transfers are right: in production, handing a caller to a person can be the right call. On this exam, every task was designed for the agent to finish, so every transfer counts as unnecessary.
  • Simulated callers: the angry persona is one model of an angry customer, not a recording of real ones.

Test your agent on your angriest callers

We found this in 3,635 simulated calls. A demo rarely shows how an agent handles a caller who pushes back. A simulation shows it with callers who are calm, angry, or somewhere in between.

See how your agent handles pushback

Veris runs your agent against simulated callers and twins of your systems, and grades what it did, including whether it kept the calls it should have kept.

Book a demo