Living definition
What is LMCache?
MLOps & AI Infrastructure Compute & Web Infra
- As of
- 2026-07-22
- Revision
- 2026-07-22.1
- Method
- v1.3.0
Definition
The term, in context
AI-assisted draft · approved dataset
LMCache is a library that plugs into large-language-model inference engines to turn the key-value cache generated while processing a prompt into a persistent, shareable resource, instead of a per-request, GPU-only structure that gets discarded. It stores cache chunks across a tiered hierarchy of GPU memory, system memory, disk, and remote storage, so repeated or overlapping prompts can reuse previously computed cache rather than recomputing it.
Live readiness status
Status as of the current dateline
AI-assisted assembly · derived results
As of 2026-07-22, verified readiness is 64 (claimed 75, reported 72, gap 11) — Critical evidence strength; current signals suggest Track; not yet.
The readiness fact belongs to the canonical Readiness Verdict for LMCache.
Dataset approval
Human editorial release
pragma.vision editorial — standing authorization (operator dev@soft.house, 2026-07-17)
- Dataset
- 2026-07-22.1
- Hash
- sha256:2bf83ba8fe86d0c871510bbef8b13d2db675e4d9efe9c64fe0e754acb6688534
- Approved
- 2026-07-22