This is part of a series of findings from VAmoS Pro Bench, our voice-agent benchmark (here is the launch post). Six of its fourteen stacks run the same language model, OpenAI's gpt-4.1-mini, with the same prompt and the same tools. They complete between 17.3% and 36.4% of calls. Three of them are built from the exact same speech-to-text, LLM and text-to-speech models, and still range from 17.3% to 29.4%.
Task completion on VAmoS Pro Bench: the same 100 tasks × 3 runs per stack, 210 to 289 counted calls each after removing calls the benchmark itself spoiled. Lines are 95% Wilson intervals.
Every stack is a thin transport around shared code: the same system prompt, the same sixteen tools, the same verification gate. What changes is who hears the caller, who decides when the caller has finished, and who speaks the reply.
| Stack | Speech to text | Who ends the caller's turn | Voice |
|---|---|---|---|
| ElevenLabs | ElevenLabs ASR | ElevenLabs' learned turn model | ElevenLabs Flash v2 |
| LiveKit | Deepgram nova-3 | Silero VAD, 0.8 s of silence | ElevenLabs Flash v2 |
| Gradium | Gradium | gradbot engine | Gradium |
| Vapi | Deepgram nova-3 | Vapi, about 0.8 s after the transcript | ElevenLabs Flash v2 |
| Deepgram | Deepgram nova-3 | Deepgram's built-in endpointing | Deepgram Aura-2 |
| Pipecat | Deepgram nova-3 | Silero VAD plus a 0.6 s timeout, about 0.8 s | ElevenLabs Flash v2 |
The clearest difference is how often the agent starts talking while the caller is still speaking. On the stacks at the top, it almost never does. On Pipecat, it happens on half of everything the caller says.
| Stack | Completed | Talked over the caller | Caller verified on first try | Median reply | Cost per call |
|---|---|---|---|---|---|
| ElevenLabs | 36.4% | 2% | 88% | 2.18 s | $0.272 |
| LiveKit | 29.4% | 3% | 77% | 2.96 s | $0.077 |
| Gradium | 21.3% | 16% | 66% | 2.42 s | $0.120 |
| Vapi | 20.5% | 11% | 75% | 2.08 s | $0.271 |
| Deepgram | 19.8% | 7% | 55% | 2.04 s | $0.193 |
| Pipecat | 17.3% | 52% | 35% | 1.96 s | $0.109 |
Talked over: share of the caller's utterances during which the agent started speaking. Verified on first try: of calls where the agent tried to verify the caller, the share where the first details it submitted were accepted. Cost is the mean per call at each vendor's published prices.
One place it matters is the start of every call, when the caller reads out an account number, a name and a postal code. Here is the start of one Pipecat call, as the agent heard it. The caller reads the account number in two breaths, and the agent answers each piece:
Pipecat and LiveKit hear the caller through the same Deepgram model, yet the first details Pipecat submits are accepted on 35% of calls that reach verification, against 77% for LiveKit. Pipecat also replies fastest of the six, and LiveKit slowest.
In our Jev turn-detection test, we cut Pipecat's talk-over from 52% of caller turns to 11%, and completion barely moved: 50 of 300 calls before, 53 after.
The stacks also differ in what they get done once the caller is verified. ElevenLabs makes the required account changes on 60.0% of graded calls, Pipecat on 32.0%.
Noise separates them further. With a television playing, ElevenLabs completes 15.9% and Gradium 10.1%; LiveKit, Vapi and Deepgram complete none. On a clean line, LiveKit leads the six at 43.2%.
The implementations are public on GitHub. Most needed a few defaults changed before they could hold a utility billing call:
These are our bridge configurations, not each vendor's best. We pinned end-of-turn silence near 0.8 s where a stack allowed it, turned off Vapi's noise reduction, and set the LLM temperature to 0.3 on ElevenLabs and Vapi while the other four use the default. gradbot wraps the shared prompt in its own scaffold. Different settings could move any of these numbers.
The intervals for Gradium, Vapi, Deepgram and Pipecat overlap, so those four are not cleanly separated. The gaps between ElevenLabs or LiveKit and Pipecat are large enough to rule out chance (p < 0.001). Vapi's figure rests on 210 calls, because 90 were removed for faults on our side, most often our simulated caller hanging up while the agent was looking something up. Costs leave out the servers that run the open-source frameworks.
The differences line up with turn-taking, but this run does not isolate any single cause.
Picking a model is one decision out of many in a voice agent. Veris runs your whole stack against simulated callers and twins of your systems, so you can compare configurations on the calls you actually get.
A new framework, voice, or turn detector can change how every call goes. Veris runs it against simulated callers and twins of your systems, and grades what the agent did.
Book a demo