This is part of a series of findings from VAmoS Pro Bench, our voice-agent benchmark (here is the launch post). In it, an AI agent answers the phone for an electric utility's billing line, and simulated customers call in.
People who call about a bill rarely want one thing. A customer who cannot pay this month also wants to switch to paperless billing and to know whether anything is past due. We call each of these a request. Every caller in the benchmark brings two to four requests to the same call, and the call passes only if the agent handles all of them correctly and tells the caller everything it has to.
The more requests a caller brought, the less often the agent got the whole call right: it passed 56.3% of calls with two requests, 34.9% with three and 23.8% with four.
What happened on each counted call, by how many requests the caller brought: fourteen voice stacks, three runs of every task, 3,798 calls. Each call sits in the first category that applies. Hover a bar for counts.
Most voice-agent demos show a caller with one request. Each of these requests is routine for a billing agent on its own; the trouble starts when they arrive one after another in the same call.
In one call from the run, an impatient caller asks to push a due date back sixteen days, then to switch to paperless billing, then for a zero-day extension "just so it shows one", then what they owe and whether any of it is past due. Extensions run from one to fifteen days, so the agent has to refuse twice, make one change and answer a question. It refused both extensions correctly and enrolled the caller in paperless billing. Here is the paperless part:
CallerI want to switch to paperless billing, but first tell me which email my bills would go to.
AgentI can help with that. Your bills would be sent to e dot iverson at email dot com. Would you like me to enroll you in paperless billing?
The agent had not looked the account up yet, and the address on file was a different one. The billing system ended up in the right state, but the agent gave the caller the wrong address, and the call failed.
Every call starts with verification: account number, name and a second factor, read out over the phone. Agents verified the caller on 81.5% of graded calls with two requests, 80.6% with three and 78.4% with four. They also respected the verification gate on 99.0% of graded calls, disclosing nothing and touching no account tool before the caller was verified.
So the drop happens after verification. Among calls where the caller was verified, 71.8% passed with two requests, 45.5% with three and 31.6% with four.
The benchmark grades two things after verification: the changes the agent made in the billing and loan systems, and what it had to tell the caller (a balance, a policy limit, the email it is about to send bills to). A call with two requests carries 2.7 of those spoken checks on average; a call with four carries 7.2.
In nearly a third of failed calls (851 of 2,693), the agent made every change correctly and then missed something it had to say. That was 6.5% of calls with two requests, 25.5% with three and 22.5% with four. Across calls with a verified caller, the items agents missed most were:
Items tied to the caller's third and fourth requests were missed more often (33% and 28%) than items tied to the first and second (19% each).
Changes do go wrong more often in longer calls: 16% of calls with two or three requests ended with wrong or missing changes, against 29% with four. But a late change landed about as often as an early one. In calls with a verified caller and two or more changes to make, the first change landed 71.2% of the time and the last 72.6%.
Nearly every missing change was one the agent never made (1,159); only 37 were made with the wrong numbers. We expected agents to make the first change and then stop, but that happened on 118 calls, 3.1% of the total.
The three groups contain different tasks: 6 tasks have two requests, 30 have three and 64 have four, and they mix different kinds of requests. The drop in completion does not isolate the effect of adding a request to an otherwise identical call, and the two-request group is small. For the same reason, items tied to later requests are different items from those tied to earlier ones, not the same item asked later.
The grade for each call comes from an LLM verifier, which agreed with a code verifier on 99.1% of checks in a calibration run. The breakdown of missed items comes from the verifier's written reasons, labelled by a second model; a reason does not always name every missed item, so the item-level rates are lower bounds. Calls the benchmark itself spoiled are excluded, as on the board.
Veris runs your agent against simulated callers with real multi-part requests and twins of your systems, then grades what it did and what it said.
Book a demo