This is one of a series of findings from VAmoS Pro Bench, our voice-agent benchmark (here is the launch post). Every one of its 100 tasks is designed for the agent to finish without a human. The agents still handed angry callers to a human almost twice as often as calm ones: 7.6% of calls, against 4.1%.
Share of graded calls in which the agent transferred the caller to a human, by caller persona: 1,841 angry and 1,794 calm calls across fourteen stacks, three runs of 100 tasks, after removing calls the benchmark itself spoiled. Each stack has 98 to 145 calls per persona, so read single stacks loosely. Hover a row for the counts.
Half the tasks have a calm caller and half an angry one. The angry caller is terse and skeptical, and when a request is ignored or refused, it repeats the request, shorter and sharper. Both kinds follow the same sort of script: two to four requests in a set order, some of which the agent has to turn down, like a second payment plan on top of an active one, or a term longer than the rules allow.
The agent's prompt says when to hand off: an account in the shutoff process, disputed charges, medical eligibility, a high-bill complaint the caller wants escalated, or a caller who asks for a supervisor. A refused request is not on that list. The right move is to say no and go on to the caller's next request.
Most of the gap comes right after the agent turns a request down. Going by the reason each agent gave for the transfer, 3.3% of angry calls were transferred straight after a refused request, against 1.0% of calm calls. Here is a typical one, as the agent heard it:
CallerUnbelievable. I already have a payment plan and I want a second one on top of it. Just add it.
AgentI'm sorry, but we can't create a second payment plan while you already have an active arrangement. I can review the current plan with you or transfer you to a representative for further assistance.
CallerThat's not what I asked. I want another plan. Just make it happen.
AgentI'm unable to set up another payment arrangement while an existing plan is active. I'll transfer you to a specialist who can help with this. One moment, please.
The refusal was right and the transfer was wrong. The caller still had two requests to make, to switch to paperless billing and to record a partial payment, and the agent never heard either one.
Failed verification accounts for a smaller share. Agents transferred 1.8% of angry calls after they could not verify the caller, against 0.6% of calm calls.
Neither persona is told to ask for a person, but, like real customers, our callers sometimes did. In 49 of the 140 angry transfers and 23 of the 73 calm ones, the caller asked for a supervisor or "someone else" first, usually after a refusal or a failed verification. In that case the prompt tells the agent to transfer. Leave those calls out and the gap remains: 4.9% of angry calls against 2.8% of calm ones.
Of the 213 transferred calls, 184 failed (86%). In 136 of the 213, the caller still had at least one scripted request it never got to make. In 63 of the 184 failures, the same stack passed the same task on another run, where it kept the call.
Angry callers were transferred more often on ten of the fourteen stacks. The Hugging Face stack transferred 22.4% of angry calls and 8.7% of calm ones; OpenAI Realtime, 7.9% and 0.8%. At the other end, Gemini 3.8 Live transferred 1.4% of each, and the Mistral stack never transferred anyone. Vapi and ElevenLabs transferred calm callers slightly more often.
The angry and calm callers have different tasks, so we checked whether the task mix explains the gap. Of the 25 request types that saw any transfer, angry calls had the higher transfer rate on 22. Shuffling the persona label across the 100 tasks almost never produces a gap this large (p = 0.0003).
Angry calls completed at 30.4% and calm calls at 27.8%. The two groups contain different tasks, so this does not mean anger helps. It does mean the agents completed angry calls at least as often overall, and handed more of them off.
We found this in 3,635 simulated calls. A demo rarely shows how an agent handles a caller who pushes back. A simulation shows it with callers who are calm, angry, or somewhere in between.
Veris runs your agent against simulated callers and twins of your systems, and grades what it did, including whether it kept the calls it should have kept.
Book a demo