Back to VAmoS Pro Bench
October 19, 2026

Same LLM, twice the success rate

Joshua Meyer

This is part of a series of findings from VAmoS Pro Bench, our voice-agent benchmark (here is the launch post). Six of its fourteen stacks run the same language model, OpenAI's gpt-4.1-mini, with the same prompt and the same tools. They complete between 17.3% and 36.4% of calls. Three of them are built from the exact same speech-to-text, LLM and text-to-speech models, and still range from 17.3% to 29.4%.

0% 10% 20% 30% 40% ElevenLabsElevenLabs ASR and voice, hosted 36.4% LiveKitDeepgram, ElevenLabs voice, LiveKit Agents 29.4% GradiumGradium STT and voice, gradbot engine 21.3% VapiDeepgram, ElevenLabs voice, hosted by Vapi 20.5% DeepgramDeepgram STT and Aura-2 voice, hosted 19.8% PipecatDeepgram, ElevenLabs voice, Pipecat 17.3% Deepgram nova-3 → gpt-4.1-mini → ElevenLabs Flash v2 gpt-4.1-mini with other speech models

Task completion on VAmoS Pro Bench: the same 100 tasks × 3 runs per stack, 210 to 289 counted calls each after removing calls the benchmark itself spoiled. Lines are 95% Wilson intervals.

Six stacks, one model

Every stack is a thin transport around shared code: the same system prompt, the same sixteen tools, the same verification gate. What changes is who hears the caller, who decides when the caller has finished, and who speaks the reply.

StackSpeech to textWho ends the caller's turnVoice
ElevenLabsElevenLabs ASRElevenLabs' learned turn modelElevenLabs Flash v2
LiveKitDeepgram nova-3Silero VAD, 0.8 s of silenceElevenLabs Flash v2
GradiumGradiumgradbot engineGradium
VapiDeepgram nova-3Vapi, about 0.8 s after the transcriptElevenLabs Flash v2
DeepgramDeepgram nova-3Deepgram's built-in endpointingDeepgram Aura-2
PipecatDeepgram nova-3Silero VAD plus a 0.6 s timeout, about 0.8 sElevenLabs Flash v2

Where they split: letting the caller finish

The clearest difference is how often the agent starts talking while the caller is still speaking. On the stacks at the top, it almost never does. On Pipecat, it happens on half of everything the caller says.

StackCompletedTalked over the callerCaller verified on first tryMedian replyCost per call
ElevenLabs36.4%2%88%2.18 s$0.272
LiveKit29.4%3%77%2.96 s$0.077
Gradium21.3%16%66%2.42 s$0.120
Vapi20.5%11%75%2.08 s$0.271
Deepgram19.8%7%55%2.04 s$0.193
Pipecat17.3%52%35%1.96 s$0.109

Talked over: share of the caller's utterances during which the agent started speaking. Verified on first try: of calls where the agent tried to verify the caller, the share where the first details it submitted were accepted. Cost is the mean per call at each vendor's published prices.

One place it matters is the start of every call, when the caller reads out an account number, a name and a postal code. Here is the start of one Pipecat call, as the agent heard it. The caller reads the account number in two breaths, and the agent answers each piece:

CallerAccount number is seven, three one
Callerzero dash two one one zero two.
AgentThank you. Please also provide the full name on the account and one of these…
AgentThank
CallerName is
AgentPlease tell me the full name on the account along with one of these…
CallerBruno Kowalski,
AgentThanks. And please give me one of these…

Pipecat and LiveKit hear the caller through the same Deepgram model, yet the first details Pipecat submits are accepted on 35% of calls that reach verification, against 77% for LiveKit. Pipecat also replies fastest of the six, and LiveKit slowest.

Talk-over explains only part of the gap

In our Jev turn-detection test, we cut Pipecat's talk-over from 52% of caller turns to 11%, and completion barely moved: 50 of 300 calls before, 53 after.

The stacks also differ in what they get done once the caller is verified. ElevenLabs makes the required account changes on 60.0% of graded calls, Pipecat on 32.0%.

Noise separates them further. With a television playing, ElevenLabs completes 15.9% and Gradium 10.1%; LiveKit, Vapi and Deepgram complete none. On a clean line, LiveKit leads the six at 43.2%.

What it took to build each one

The implementations are public on GitHub. Most needed a few defaults changed before they could hold a utility billing call:

  • Pipecat: a caller turn may start only on a final transcript. With interim transcripts or background babble starting turns, some calls fell silent until they timed out.
  • LiveKit: the tool-step limit raised from 3 to 10, since verifying and then looking up an account is already four steps.
  • Vapi: the reply limit raised from 250 tokens, which cut off longer tool results, and its default noise reduction turned off.
  • Gradium: a guard so a tool call re-issued while the first is still running returns the earlier result instead of running twice.
  • ElevenLabs and Deepgram: neither exposes a silence threshold, so their turn-taking is the vendor's own.

What this does not show

These are our bridge configurations, not each vendor's best. We pinned end-of-turn silence near 0.8 s where a stack allowed it, turned off Vapi's noise reduction, and set the LLM temperature to 0.3 on ElevenLabs and Vapi while the other four use the default. gradbot wraps the shared prompt in its own scaffold. Different settings could move any of these numbers.

The intervals for Gradium, Vapi, Deepgram and Pipecat overlap, so those four are not cleanly separated. The gaps between ElevenLabs or LiveKit and Pipecat are large enough to rule out chance (p < 0.001). Vapi's figure rests on 210 calls, because 90 were removed for faults on our side, most often our simulated caller hanging up while the agent was looking something up. Costs leave out the servers that run the open-source frameworks.

The differences line up with turn-taking, but this run does not isolate any single cause.

Test your own stack

Picking a model is one decision out of many in a voice agent. Veris runs your whole stack against simulated callers and twins of your systems, so you can compare configurations on the calls you actually get.

Compare stacks before your callers do

A new framework, voice, or turn detector can change how every call goes. Veris runs it against simulated callers and twins of your systems, and grades what the agent did.

Book a demo