Comparative expertise
Real-time conversational voice AI
One fixed criteria frame. Every populated cell traces to one dated readiness verdict; missing evidence remains explicit.
- Data as of
- 2026-08-20
- Definition pinned
- 2026-08-12
- Revision
- 1
- Definition version
- v1.0.0
- Method
- v1.3.0
Fixed matrix
Verdict fields, side by side
AI-assisted assembly · derived results
Scope: Voice & Speech AI · Agentic Commerce & Payments. No average, blend, composite, or estimate is produced.
| Criterion | ElevenLabs Agents elevenlabs-agents | OpenAI GPT-Realtime-1.5 openai-gpt-realtime-1-5 | Azure Realtime Voice (Build 2026) azure-realtime-voice-build-2026 | NVIDIA PersonaPlex-7B nvidia-personaplex-7b |
|---|---|---|---|---|
| Verified readiness verdict.readiness.verified | 56 | 62 | 69 | 55 |
| Hype gap verdict.readiness.gap | 34 | 23 | 16 | 15 |
| Evidence strength verdict.readiness.strength | Critical | Critical | Strong | Critical |
| Recommended stance verdict.decision.answer | Proceed with caution | Adopt with guardrails | Adopt with guardrails | Track; not yet |
| Latest evidence verdict.evidence_summary.latest_observed_on | 2026-06-25 | 2026-06-25 | 2026-06-05 | 2026-06-25 |
Human editorial
Synthesis and caveats
Human editorial · 2026-08-12
Synthesis
The four systems occupy different layers of the real-time voice stack, which is itself the most load-bearing fact for reading this cluster. ElevenLabs Agents -- rebranded ElevenAgents on 2026-02-09 -- is a managed, full-stack orchestration platform (speech recognition, interchangeable LLM, synthesis, turn detection, telephony, and monitoring bundled together), not a single model; Revolut's named production deployment (announced 2026-01-28) put it in front of more than four million customers across 30+ languages with a reported 99.7% call-success rate, and OpenBenchmarks' PSTN latency test (last run 2026-08-01) measured its p95 tail latency as the best of five tested platforms even though its median trailed Telnyx. OpenAI's GPT-Realtime-1.5 is a hosted speech-to-speech model exposed through the Realtime API rather than a complete agent platform; it is also the most dated of the four -- released 2025-02-23 and already superseded twice, by GPT-Realtime-2 (2025-05-07) and GPT-Realtime-2.1 (2026-07-06) -- and independent evaluation is mixed: Scale Labs' blind Voice Showdown (published 2026-03-20) found it lost roughly three-quarters of head-to-head battles against its own predecessor and fell below a 50% win rate in every tested non-English language, while a separate reproducible tool-calling study (Laskar et al., revised 2026-05-20) had it leading the When2Call benchmark at 71.9. Microsoft's entry conflates two distinct objects under one register row: the azure-realtime speech-to-speech model and the Voice Live API that wraps it into an agent platform comparable to ElevenAgents; Microsoft's own release notes describe azure-realtime as production-ready, but an independent second-source check could not locate a standalone GA announcement confirming the specific date, so treat GA timing as vendor-reported and unconfirmed, and note that Voice Live's hosted-agent and WebRTC transport paths remain in public preview even where the underlying model is not. NVIDIA's PersonaPlex-7B is the clear outlier: an open-weights, MIT/NVIDIA Open Model License research checkpoint released 2026-01-15, self-hosted on customer hardware rather than offered as a managed API, with no named production customer -- and NVIDIA's own newer Nemotron 3 VoiceChat (announced 2026-03-24) already positions PersonaPlex as the research predecessor rather than the production path forward. Independent reproduction of PersonaPlex's own published benchmarks found materially worse turn-taking latency than NVIDIA's reported figures (Ohashi et al., 2026-06-09), and a separate interruption-robustness study found its response correctness collapsed from 49% to 28% under abrupt interruption (Chang et al., 2026-06-09) -- exactly the kind of failure mode aggregate demo videos do not surface. Read together: a buyer choosing a deployable, supported voice-agent platform today is choosing between ElevenAgents and Azure Voice Live, not between all four; a developer choosing a raw hosted speech model is choosing between GPT-Realtime-1.5 and azure-realtime; and PersonaPlex belongs in a different conversation entirely -- open-weights research versus managed product -- despite sharing this register's discipline label with the other three.
Caveats
This cluster deliberately mixes object types (a managed platform, two hosted model APIs, and one open-weights research checkpoint) that the register groups under one discipline; treat it as a landscape read, not a like-for-like benchmark. Azure's azure-realtime GA date is vendor-stated and not independently corroborated by a second primary source as of this reading -- re-verify before citing a specific date. ElevenLabs' Revolut figures (call-success rate, resolution-time improvement) are a vendor-hosted customer case study, not an independently audited report. GPT-Realtime-1.5's Scale Labs and Laskar et al. results predate this reading by several months and describe the 1.5 generation specifically -- they say nothing about GPT-Realtime-2/2.1's current behavior. PersonaPlex's independent reproduction studies (Ohashi et al.; Chang et al.) are both non-peer-reviewed preprints as of 2026-06-09.
Approved by pragma.vision editorial — standing authorization (operator dev@soft.house, 2026-07-17) · 2026-08-12