pragma.vision Technology observatory

Verification register Compute & Web Infra

Living definition

What is Mooncake?

MLOps & AI Infrastructure Compute & Web Infra

As of
2026-07-22
Revision
2026-07-22.1
Method
v1.3.0

Definition

The term, in context

AI-assisted draft · approved dataset

Mooncake is a large-language-model serving architecture that treats the key-value cache, rather than the GPU itself, as the primary resource to schedule. It separates prompt processing and token generation onto independent GPU clusters connected over a network, pooling their spare memory and storage capacity into a shared cache store that both clusters read from and write to.

Live readiness status

Status as of the current dateline

AI-assisted assembly · derived results

As of 2026-07-22, verified readiness is 71 (claimed 80, reported 77, gap 9) — Strong evidence strength; current signals suggest Too early to adopt.

The readiness fact belongs to the canonical Readiness Verdict for Mooncake.

Dataset approval

Human editorial release

pragma.vision editorial — standing authorization (operator dev@soft.house, 2026-07-17)

Dataset
2026-07-22.1
Hash
sha256:2bf83ba8fe86d0c871510bbef8b13d2db675e4d9efe9c64fe0e754acb6688534
Approved
2026-07-22

Your opinion

Tell us anything.

What works, what doesn't, what's missing — especially about our watches, lenses, and the register itself. Anonymous is fine; leave an email if you'd like a reply.