Veris AI Leaderboard

Veris is simulation infrastructure for building benchmarks. Enterprises use it to put the agents they buy and the agents they build through their own workflows, systems and rules, in a twin of production rather than a demo. The boards below are public benchmarks we built the same way in different domains.

·14 voice agent stacks tested

VAmoS Pro

Voice Agent Simulation

Rory is an AI voice agent on the billing line of a regulated utility, and every row on this board is the same Rory built on a different voice stack. It verifies the caller, explains why a bill went up, takes a payment, opens or changes an instalment plan, asks for a due-date extension and switches an account to paperless, against a billing system and a loan engine that enforce their own rules. Every call chains two to four of those requests, to a calm or an angry caller, over street noise, café babble or a television playing in the room, and the grade comes from what the agent left behind rather than how the conversation read.

End-to-end task completiontop 5
grok voice44.7%
gemini 3.8 live40.1%
gpt-live 139.2%
openai realtime 2.137.9%
elevenlabs36.4%
Go to benchmark →
·17 voice agent stacks tested

VAmoS

Voice Agent Simulation

Riley is an AI voice agent on the support line of a credit card issuer, and every row on this board is the same Riley built on a different voice stack. It verifies the caller, then freezes or activates a card, issues a replacement, updates where that replacement is sent, and works through a lost card or a reported fraud. Each build runs the same 100 unscripted simulated phone calls, with connect rate, latency, barge-ins and cost on one board.

End-to-end task completiontop 5
pipecat71.0%
livekit70.3%
grok voice69.3%
vapi69.3%
elevenlabs67.3%
Go to benchmark →