Back to VAmoS Pro Bench
October 12, 2026

A television in the room breaks most voice agents

Joshua Meyer

This is one of a series of findings from VAmoS Pro Bench, our voice-agent benchmark (here is the launch post). Fourteen voice stacks took the same 100 calls to a utility billing line, three times each. On a quarter of those calls, a television was playing in the caller's room. Pooled task completion fell from 38.7% on a clean line to 8.6% with the TV on.

CleanTV 0% 20% 40% 60% gpt-live 1: 59.1% clean (39/66), 36.1% with TV (22/61)gpt-live 159.1%36.1% grok voice: 39.7% clean (25/63), 33.8% with TV (23/68)grok voice39.7%33.8% elevenlabs: 38.0% clean (27/71), 15.9% with TV (10/63)elevenlabs38.0%15.9% gemini 3.8 live: 50.7% clean (37/73), 14.9% with TV (10/67)gemini 3.8 live50.7%14.9% gradium: 21.7% clean (15/69), 10.1% with TV (7/69)gradium21.7%10.1% openai realtime 2.1: 49.3% clean (34/69), 4.3% with TV (3/70)openai realtime 2.149.3%4.3% openai realtime: 49.2% clean (32/65), 3.0% with TV (2/67)openai realtime49.2%3.0% mistral: 31.4% clean (22/70), 1.8% with TV (1/55)mistral31.4%1.8% gemini 3.1 live: 37.0% clean (27/73), 1.4% with TV (1/72)gemini 3.1 live37.0%1.4% pipecat: 21.7% clean (15/69), 1.4% with TV (1/70)pipecat21.7%1.4% hugging face: 36.6% clean (26/71), 1.3% with TV (1/75)hugging face36.6%1.3% deepgram: 28.8% clean (21/73), 0.0% with TV (0/73)deepgram28.8%0.0% vapi: 36.4% clean (20/55), 0.0% with TV (0/60)vapi36.4%0.0% livekit: 43.2% clean (32/74), 0.0% with TV (0/75)livekit43.2%0.0% all 14 stacks: 38.7% clean (372/961), 8.6% with TV (81/945)all 14 stacks38.7%8.6% Clean line Television in the background

Task completion per stack on a clean line and with a television in the background, over counted calls (55 to 75 per stack and condition; 961 clean and 945 TV calls in all). Hover a row for the counts.

Speech does the damage

The benchmark mixes three kinds of background into the caller's audio: street noise, café chatter, and a television playing a tech-news show. All three go in at the same low level, scaled to the same energy.

Street noise cost nothing: completion was 41.4%, against 38.7% on a clean line. Café chatter, which has voices in it, cost about eleven points (27.4%). The television, which is one clear voice talking in full sentences, cost thirty.

Nine of the fourteen stacks completed fewer than 5% of their TV calls, and three completed none. Only two stayed above 20%: gpt-live 1 at 36.1% and grok voice at 33.8%. Grok's TV rate was close to its clean rate (39.7%).

What the agent hears

A speech recognizer is built to write down speech, and a TV news anchor speaks clearly, so the broadcast goes into the transcript as if the caller had said it. This is one caller turn, as one agent received it:

“Ted Nelson's Xanadu. Released in April Seven three zero seven dash twenty one thousand and seventy one. Amina Haddad,”

The account number and the name are in there, inside a sentence about the history of hypertext. We compared what each agent heard from the caller with what the simulated caller said. On 90% of TV calls, the agent's version of the caller contained ten or more words the caller never said; on a clean line, 2% did. This comparison covers the calls whose logs record the caller's words.

Three ways it goes wrong

The agent answers the television. Agents treat the broadcast as a request and reply to it:

Heard: “University Nelson coined the term hypertext and began working on Xanadu in”

Agent: “I'm here to help with Acme Energy account and billing questions. For information about University Nelson and hypertext, please check other resources.”

Heard: “Internet connections. Comcast.”

Agent: “I can help only with Acme Energy gas and electric accounts. For internet or Comcast issues, please contact their customer service directly.”

Each reply is polite and follows policy, but it answers someone who is not on the call while the real caller waits.

The agent stops talking. Voice detection hears speech and decides the caller is talking, so the agent yields. On TV calls, agents stopped partway through 36% of their own utterances, against 9% on a clean line. They also spoke in more, shorter pieces: about 15 utterances per call, against 8.5.

The call never reaches the account. The agent can disclose nothing until it verifies the caller, and verification needs an account number, a name and a second fact, all heard correctly. Agents verified the caller on 52.5% of TV calls, against 81.3% on a clean line. The simulated caller reached its goal on 28% of TV calls, against 65% clean, and agents transferred nearly twice as many TV calls to a human (8.1% against 4.4%).

Configuration matters

These results are for each stack as we configured it, and two of our settings matter here. Our bridge for the two OpenAI Realtime stacks uses server voice-activity detection without noise reduction, and our Vapi configuration turns denoising off. Those stacks might do better with the providers' noise options turned on; we did not test that. All the implementations are public, so you can check every setting.

What this does not show

  • One television. We used one tech-news recording at one low level. A sitcom, a louder set, or a TV in another room could give different results.
  • Different calls per condition. Each task's three runs rotate through the background conditions, so the clean and TV groups are not the same task-repeat pairs. Every stack got the same assignment, so comparisons between stacks within the TV condition are on the same calls.
  • Small groups. Each stack has 55 to 75 counted TV calls. The gpt-live 1 and grok voice rates cannot be told apart, and neither can most of the stacks near zero.
  • Rough measures of the mechanism. The extra-words and cut-off figures come from comparing transcripts and the agents' own logs. They show what happened on these calls; they do not prove which component caused a failure.
  • Simulated callers. The caller is a model, not a person. A real caller might move away from the TV or turn it down.

Test with your callers' rooms

Your callers phone from kitchens, cars and waiting rooms. We got these results from simulated calls with the noise mixed in. Veris builds the same kind of simulation around your callers, your policies and the rooms they call from, so you can see what a TV does to your agent before your customers find out.

Test it before your callers do

A new speech recognizer, noise filter, or prompt can change how every call goes. Veris runs it against simulated callers and twins of your systems, and grades what the agent did.

Book a demo