This is one of a series of findings from VAmoS Pro Bench, our voice-agent benchmark (here is the launch post). On 2,072 of its calls, the simulated caller hung up satisfied: everything it had called about was done, as far as it could tell. Of those, 1,122 (54%) failed the grade, and on 412 the agent had left the account wrong.
Every VAmoS Pro Bench call that the simulated caller ended with its goals met, after removing calls the benchmark itself spoiled: 2,072 calls, 80 to 195 per stack. The labels on the right give each stack's failed share. Stacks reach a satisfied ending on different tasks, so the bars do not compare stacks on the same calls.
What "satisfied" means here
Each VAmoS Pro call has a simulated caller with two to four things it needs done on its utility account, such as recording a payment, opening a payment plan, or switching to paperless billing. It can give up, get transferred, or decide that its goals were met and hang up.
"Goals met" is the caller's own judgment, made from what it heard. It is the closest a benchmark gets to a customer who says thanks and hangs up happy, the same signal that post-call surveys and end-of-call sentiment scores measure.
The caller cannot see the account. It hears "your payment is recorded", but it cannot see whether a tool recorded it, or for how much. The grade can. It checks the changes the agent made in the billing and loan systems behind it, and every figure and fact the agent was required to say.
Two calls that ended happy
In one call, the caller asked to record an installment it had paid that day, then to switch to paperless billing. The agent turned paperless on. Then:
Caller: Okay, finally. So we're all set, yeah?
Agent: Yes, your payment is recorded and paperless billing is set.
The agent never called the tool that records a payment. The caller hung up satisfied.
In another, the caller wanted a five-month payment plan and a $25 partial payment recorded against it. The agent tried to record the payment before the plan existed, and the loan system replied that there was no active plan. The agent then opened the plan and said:
Agent: I've set up the 5 month plan for you, and recorded that twenty five dollar payment against it.
Caller: Yes, fine. I agree. Do it. And that's all.
The plan was real. The agent never tried the payment again, so it was never recorded.
What goes wrong behind a happy ending
We split the 1,122 failed calls by what the grade found. A call can fail more than one check, so the counts within each group overlap.
The account was left wrong: 412 calls, or 20% of all satisfied endings.
A requested change never happened (313 calls). In these calls, paperless billing was never turned on (216), the payment was never recorded (88), or the plan was never opened (53). In 264 of the 313, the agent never called the tool at all.
The right changes in the wrong order (48). The benchmark checks the order of the changes the task asks for.
An extra or repeated change (39), such as recording the same payment twice.
The right change with the wrong numbers (12).
The account was right, but the agent left out or got wrong something it had to say: 710 calls, or 34%.
A figure the caller needed (424 calls), such as the balance, the installment amount, the payment, or the next due date.
A fact about the account (334), such as whether it was past due, whether a plan was already active, or whether the one due-date extension had already been used.
The policy limit (120): the longest or shortest term the account allows.
Consent (84): the agent turned on paperless billing without reading back the email address and waiting for a yes (72), or opened a plan without quoting it first and waiting for agreement (12).
Every stack does this: between 46% and 70% of each stack's satisfied calls failed. The share where the account was left wrong ranges from 8% for gpt-live 1 to 37% for pipecat.
The signal also fails the other way. Of the 1,105 passing calls, 155 (14%) did not end happy. The caller hung up after the agent turned down a request (67), gave up (63), or was transferred (23). Turning a caller down is sometimes the right answer under the utility's policy, and callers do not always take it well.
What to measure instead
If you score a voice agent on whether callers ended the call happy, by survey, end-of-call sentiment or a simulated caller's own verdict, you are scoring what the agent said. On this exam, more than half of the calls that ended happy had failed, and one in five had left the customer's account wrong.
Grade the world instead: check what the agent changed in your systems against what the caller asked for and what your policy allows, and check what it told the caller against the record.
What this does not show
The caller is simulated. "Goals met" is a language model's judgment from what it heard, not a survey of real customers.
The grade is strict by design. It requires the exact changes, in order, and specific statements. A few of the 710 "said" failures are close calls, such as an agent saying "two installments" where the check wanted "two months". We removed calls whose verdicts an audit found wrong, but unaudited verdicts can still be wrong. For the 412 calls where the account was left wrong, we rechecked each one directly against the tool calls in its trace.
Stacks are compared on different calls. Each stack reaches a satisfied ending on a different set of tasks, and there are only 80 to 195 such calls per stack.
This is one exam, on one utility's policies.
Grade what your agent did
A happy caller is not a finished task. Veris runs your agent against simulated callers and twins of your systems, and grades the changes it made and what it said, not how the call ended.