pragma.vision Technology observatory & strategic foresight

Verification register Compute & Web Infra

Comparative expertise

LLM inference serving engines

One fixed criteria frame. Every populated cell traces to one dated readiness verdict; missing evidence remains explicit.

Data as of
2026-08-20
Definition pinned
2026-07-17
Revision
2
Definition version
v1.0.0
Method
v1.3.0

Fixed matrix

Verdict fields, side by side

AI-assisted assembly · derived results

Scope: AI Inference & Accelerated Compute · Compute & Web Infra. No average, blend, composite, or estimate is produced.

Comparison of SGLang (v0.5.x), vLLM V1
Criterion SGLang (v0.5.x) sglang-inference-engine vLLM V1 vllm-v1-engine
Verified readiness verdict.readiness.verified 85 64
Hype gap verdict.readiness.gap 5 24
Evidence strength verdict.readiness.strength Strong Strong
Recommended stance verdict.decision.answer Adopt with guardrails Low-friction candidate
Latest evidence verdict.evidence_summary.latest_observed_on 2026-06-13 2026-06-01

Human editorial

Synthesis and caveats

Human editorial · 2026-07-17

Synthesis

Both engines are GA and iterating quickly: SGLang's v0.5.15.post1 landed 2026-07-14 (the project spun out commercially as RadixArk in January 2026), and vLLM's v0.25.1 landed the same day, with the V1 engine as default since v0.8.0 (2025-03-18). Aggregated third-party benchmarks — compiled by consultancy posts from Spheron, Prem AI, Runpod and others, not independently measured — suggest SGLang throughput advantages on H100 and DeepSeek-class workloads, while vLLM retains the broadest ecosystem, hardware support, and default status in most cloud serving stacks. A peer-reviewed ISPASS 2026 study of confidential-GPU serving splits the verdict: SGLang led at compute-bound large batches; vLLM showed lower confidential-computing overhead and better small-batch behavior. Practitioner reports name observability and governance fragmentation, operational complexity, and capacity/cost control as the production blockers. The matrix mirrors this: vLLM's 24-point claimed-versus-verified gap against SGLang's 5-point gap suggests the deployment choice is decided by workload shape — batch size, model family, confidential-compute requirements — rather than by a single winner.

Caveats

Most published head-to-heads aggregate vendor-reported numbers; treat absolute throughput figures as indicative. NVIDIA Dynamo, a frequent third option, is tracked in our register under a different discipline (MLOps & AI Infrastructure) and is therefore outside this cluster's fixed criteria frame.

Approved by pragma.vision editorial — operator dev@soft.house · 2026-07-17

Your opinion

Tell us anything.

What works, what doesn't, what's missing — especially about our watches, lenses, and the register itself. Anonymous is fine; leave an email if you'd like a reply.