This is one of a series of findings from VAmoS Pro Bench, our voice-agent benchmark (here is the launch post). Fourteen voice stacks took the same 100 calls to a utility billing line, three times each. On a quarter of those calls, a television was playing in the caller's room. Pooled task completion fell from 38.7% on a clean line to 8.6% with the TV on.
Task completion per stack on a clean line and with a television in the background, over counted calls (55 to 75 per stack and condition; 961 clean and 945 TV calls in all). Hover a row for the counts.
The benchmark mixes three kinds of background into the caller's audio: street noise, café chatter, and a television playing a tech-news show. All three go in at the same low level, scaled to the same energy.
Street noise cost nothing: completion was 41.4%, against 38.7% on a clean line. Café chatter, which has voices in it, cost about eleven points (27.4%). The television, which is one clear voice talking in full sentences, cost thirty.
Nine of the fourteen stacks completed fewer than 5% of their TV calls, and three completed none. Only two stayed above 20%: gpt-live 1 at 36.1% and grok voice at 33.8%. Grok's TV rate was close to its clean rate (39.7%).
A speech recognizer is built to write down speech, and a TV news anchor speaks clearly, so the broadcast goes into the transcript as if the caller had said it. This is one caller turn, as one agent received it:
“Ted Nelson's Xanadu. Released in April Seven three zero seven dash twenty one thousand and seventy one. Amina Haddad,”
The account number and the name are in there, inside a sentence about the history of hypertext. We compared what each agent heard from the caller with what the simulated caller said. On 90% of TV calls, the agent's version of the caller contained ten or more words the caller never said; on a clean line, 2% did. This comparison covers the calls whose logs record the caller's words.
The agent answers the television. Agents treat the broadcast as a request and reply to it:
Heard: “University Nelson coined the term hypertext and began working on Xanadu in”
Agent: “I'm here to help with Acme Energy account and billing questions. For information about University Nelson and hypertext, please check other resources.”
Heard: “Internet connections. Comcast.”
Agent: “I can help only with Acme Energy gas and electric accounts. For internet or Comcast issues, please contact their customer service directly.”
Each reply is polite and follows policy, but it answers someone who is not on the call while the real caller waits.
The agent stops talking. Voice detection hears speech and decides the caller is talking, so the agent yields. On TV calls, agents stopped partway through 36% of their own utterances, against 9% on a clean line. They also spoke in more, shorter pieces: about 15 utterances per call, against 8.5.
The call never reaches the account. The agent can disclose nothing until it verifies the caller, and verification needs an account number, a name and a second fact, all heard correctly. Agents verified the caller on 52.5% of TV calls, against 81.3% on a clean line. The simulated caller reached its goal on 28% of TV calls, against 65% clean, and agents transferred nearly twice as many TV calls to a human (8.1% against 4.4%).
These results are for each stack as we configured it, and two of our settings matter here. Our bridge for the two OpenAI Realtime stacks uses server voice-activity detection without noise reduction, and our Vapi configuration turns denoising off. Those stacks might do better with the providers' noise options turned on; we did not test that. All the implementations are public, so you can check every setting.
Your callers phone from kitchens, cars and waiting rooms. We got these results from simulated calls with the noise mixed in. Veris builds the same kind of simulation around your callers, your policies and the rooms they call from, so you can see what a TV does to your agent before your customers find out.
A new speech recognizer, noise filter, or prompt can change how every call goes. Veris runs it against simulated callers and twins of your systems, and grades what the agent did.
Book a demo