pragma.vision Technology observatory & strategic foresight

Verification register Agentic Commerce & Payments

Comparative expertise

Real-time conversational voice AI

One fixed criteria frame. Every populated cell traces to one dated readiness verdict; missing evidence remains explicit.

Data as of
2026-08-20
Definition pinned
2026-08-12
Revision
1
Definition version
v1.0.0
Method
v1.3.0

Fixed matrix

Verdict fields, side by side

AI-assisted assembly · derived results

Scope: Voice & Speech AI · Agentic Commerce & Payments. No average, blend, composite, or estimate is produced.

Comparison of ElevenLabs Agents, OpenAI GPT-Realtime-1.5, Azure Realtime Voice (Build 2026), NVIDIA PersonaPlex-7B
Criterion ElevenLabs Agents elevenlabs-agents OpenAI GPT-Realtime-1.5 openai-gpt-realtime-1-5 Azure Realtime Voice (Build 2026) azure-realtime-voice-build-2026 NVIDIA PersonaPlex-7B nvidia-personaplex-7b
Verified readiness verdict.readiness.verified 56 62 69 55
Hype gap verdict.readiness.gap 34 23 16 15
Evidence strength verdict.readiness.strength Critical Critical Strong Critical
Recommended stance verdict.decision.answer Proceed with caution Adopt with guardrails Adopt with guardrails Track; not yet
Latest evidence verdict.evidence_summary.latest_observed_on 2026-06-25 2026-06-25 2026-06-05 2026-06-25

Human editorial

Synthesis and caveats

Human editorial · 2026-08-12

Synthesis

The four systems occupy different layers of the real-time voice stack, which is itself the most load-bearing fact for reading this cluster. ElevenLabs Agents -- rebranded ElevenAgents on 2026-02-09 -- is a managed, full-stack orchestration platform (speech recognition, interchangeable LLM, synthesis, turn detection, telephony, and monitoring bundled together), not a single model; Revolut's named production deployment (announced 2026-01-28) put it in front of more than four million customers across 30+ languages with a reported 99.7% call-success rate, and OpenBenchmarks' PSTN latency test (last run 2026-08-01) measured its p95 tail latency as the best of five tested platforms even though its median trailed Telnyx. OpenAI's GPT-Realtime-1.5 is a hosted speech-to-speech model exposed through the Realtime API rather than a complete agent platform; it is also the most dated of the four -- released 2025-02-23 and already superseded twice, by GPT-Realtime-2 (2025-05-07) and GPT-Realtime-2.1 (2026-07-06) -- and independent evaluation is mixed: Scale Labs' blind Voice Showdown (published 2026-03-20) found it lost roughly three-quarters of head-to-head battles against its own predecessor and fell below a 50% win rate in every tested non-English language, while a separate reproducible tool-calling study (Laskar et al., revised 2026-05-20) had it leading the When2Call benchmark at 71.9. Microsoft's entry conflates two distinct objects under one register row: the azure-realtime speech-to-speech model and the Voice Live API that wraps it into an agent platform comparable to ElevenAgents; Microsoft's own release notes describe azure-realtime as production-ready, but an independent second-source check could not locate a standalone GA announcement confirming the specific date, so treat GA timing as vendor-reported and unconfirmed, and note that Voice Live's hosted-agent and WebRTC transport paths remain in public preview even where the underlying model is not. NVIDIA's PersonaPlex-7B is the clear outlier: an open-weights, MIT/NVIDIA Open Model License research checkpoint released 2026-01-15, self-hosted on customer hardware rather than offered as a managed API, with no named production customer -- and NVIDIA's own newer Nemotron 3 VoiceChat (announced 2026-03-24) already positions PersonaPlex as the research predecessor rather than the production path forward. Independent reproduction of PersonaPlex's own published benchmarks found materially worse turn-taking latency than NVIDIA's reported figures (Ohashi et al., 2026-06-09), and a separate interruption-robustness study found its response correctness collapsed from 49% to 28% under abrupt interruption (Chang et al., 2026-06-09) -- exactly the kind of failure mode aggregate demo videos do not surface. Read together: a buyer choosing a deployable, supported voice-agent platform today is choosing between ElevenAgents and Azure Voice Live, not between all four; a developer choosing a raw hosted speech model is choosing between GPT-Realtime-1.5 and azure-realtime; and PersonaPlex belongs in a different conversation entirely -- open-weights research versus managed product -- despite sharing this register's discipline label with the other three.

Caveats

This cluster deliberately mixes object types (a managed platform, two hosted model APIs, and one open-weights research checkpoint) that the register groups under one discipline; treat it as a landscape read, not a like-for-like benchmark. Azure's azure-realtime GA date is vendor-stated and not independently corroborated by a second primary source as of this reading -- re-verify before citing a specific date. ElevenLabs' Revolut figures (call-success rate, resolution-time improvement) are a vendor-hosted customer case study, not an independently audited report. GPT-Realtime-1.5's Scale Labs and Laskar et al. results predate this reading by several months and describe the 1.5 generation specifically -- they say nothing about GPT-Realtime-2/2.1's current behavior. PersonaPlex's independent reproduction studies (Ohashi et al.; Chang et al.) are both non-peer-reviewed preprints as of 2026-06-09.

Approved by pragma.vision editorial — standing authorization (operator dev@soft.house, 2026-07-17) · 2026-08-12

Your opinion

Tell us anything.

What works, what doesn't, what's missing — especially about our watches, lenses, and the register itself. Anonymous is fine; leave an email if you'd like a reply.