Comparative expertise
LLM inference serving engines
One fixed criteria frame. Every populated cell traces to one dated readiness verdict; missing evidence remains explicit.
- Data as of
- 2026-08-20
- Definition pinned
- 2026-07-17
- Revision
- 2
- Definition version
- v1.0.0
- Method
- v1.3.0
Fixed matrix
Verdict fields, side by side
AI-assisted assembly · derived results
Scope: AI Inference & Accelerated Compute · Compute & Web Infra. No average, blend, composite, or estimate is produced.
| Criterion | SGLang (v0.5.x) sglang-inference-engine | vLLM V1 vllm-v1-engine |
|---|---|---|
| Verified readiness verdict.readiness.verified | 85 | 64 |
| Hype gap verdict.readiness.gap | 5 | 24 |
| Evidence strength verdict.readiness.strength | Strong | Strong |
| Recommended stance verdict.decision.answer | Adopt with guardrails | Low-friction candidate |
| Latest evidence verdict.evidence_summary.latest_observed_on | 2026-06-13 | 2026-06-01 |
Human editorial
Synthesis and caveats
Human editorial · 2026-07-17
Synthesis
Both engines are GA and iterating quickly: SGLang's v0.5.15.post1 landed 2026-07-14 (the project spun out commercially as RadixArk in January 2026), and vLLM's v0.25.1 landed the same day, with the V1 engine as default since v0.8.0 (2025-03-18). Aggregated third-party benchmarks — compiled by consultancy posts from Spheron, Prem AI, Runpod and others, not independently measured — suggest SGLang throughput advantages on H100 and DeepSeek-class workloads, while vLLM retains the broadest ecosystem, hardware support, and default status in most cloud serving stacks. A peer-reviewed ISPASS 2026 study of confidential-GPU serving splits the verdict: SGLang led at compute-bound large batches; vLLM showed lower confidential-computing overhead and better small-batch behavior. Practitioner reports name observability and governance fragmentation, operational complexity, and capacity/cost control as the production blockers. The matrix mirrors this: vLLM's 24-point claimed-versus-verified gap against SGLang's 5-point gap suggests the deployment choice is decided by workload shape — batch size, model family, confidential-compute requirements — rather than by a single winner.
Caveats
Most published head-to-heads aggregate vendor-reported numbers; treat absolute throughput figures as indicative. NVIDIA Dynamo, a frequent third option, is tracked in our register under a different discipline (MLOps & AI Infrastructure) and is therefore outside this cluster's fixed criteria frame.
Approved by pragma.vision editorial — operator dev@soft.house · 2026-07-17