AI Post Transformers Visual Companion

Mooncake for KV Cache-Centric LLM Serving

Mooncake reframes long-context LLM serving around a moving cache, not just busy GPUs. These views track why TTFT and TBT pull the system in different directions, how tiered KV placement changes routing, and where Mooncake’s scheduler wins or refuses work.

arXiv 2407.00079 Version Discussed2025-09-03 Transcript IDs2407.00079 Ecosystem Snapshot2026-06-04 Themecache-first scheduling
Prefill Wide prompt pass. Compute-heavy, TTFT-sensitive, and worth splitting away from decode once prompts get long.
Decode Token-by-token loop. Memory-bandwidth and latency pressure dominate, so stable batching matters more than raw FLOPS.
KV Tiers VRAM, CPU DRAM, and SSD become a storage ladder for prompt state. Reuse and movement are inseparable.
Admission Mooncake predicts future decode pressure, then rejects early when a request is likely to miss SLOs anyway.
Reported gains in the episode cluster around three workload regimes: roughly +20–40% on public datasets, +50–525% in reuse-heavy long-context simulations, and +75% more handled requests on replayed real workloads.
Overview

TTFT Wants One System, TBT Wants Another

Mooncake separates the work that wins the first token from the work that keeps later tokens smooth. The split only pays if the transfer tax stays smaller than the prefill you avoided.

Serving Path

Switch the deployment mode to see how interference collapses back into a shared pool when prefill and decode fight for the same accelerators.

Latency Budget

Both rows render the same request class. The active focus amplifies either the first-token path or the steady decode rhythm.

Illustrative numbers are shaped by the paper’s regime: long prompts, short outputs, strong prefix reuse, and fast interconnects.
Tiered Cache

The KV Cache Becomes a Storage System

Mooncake treats cache blocks like movable assets. Reuse length, access tier, and fetch cost decide whether a request should hydrate, recompute, migrate, or replicate.

Block Heatmap by Tier

Hover blocks to inspect where reuse sits. Blue cells are cold. Orange and red cells are the blocks the scheduler wants to keep nearest to decode.

Movement Across VRAM, DRAM, and SSD

The arrows show where bytes move for the selected scenario. Wider links mean more of the request state takes that path before decode can start.

Dashed links represent async or staged hydration. Solid links are directly on the critical path to the decode pool.
Control Plane

Cache-First Routing, Then Overload-Aware Admission

Mooncake does not simply pick the emptiest GPU. It scores candidate prefill/decode pairs by reusable prefix length, remote transfer cost, queue delay, and whether decode will still hit token latency after admission.

Request Walkthrough

Step through the scheduler path. The winning route changes when the shared prefix is too short to justify remote fetch, or when decode headroom is already gone.

Why Early Rejection Is Predictive

Reactive admission waits until decode is visibly full. Predictive admission closes earlier, which avoids the queue oscillation Mooncake argues against.

Results

Biggest Wins in Reuse-Heavy, Long-Context Regimes

The paper’s strongest evidence is not “faster everywhere.” It is “much better when the workload is long, prefix-heavy, and overloaded enough that cache placement and admission actually matter.”

Workload-Specific Gain Profile

Public datasets move modestly. Synthetic long-context reuse can explode upward. Real trace replay lands between those poles and shows the architecture surviving deeper overload.

Goodput Under Offered Load

The key curve is successful requests that still meet latency objectives. Mooncake’s line stays higher for longer because it routes by cached state and rejects earlier when decode would collapse.

Bars use reported anchor ranges where the episode names them directly. Intermediate values are illustrative, meant to show mechanism and workload shape rather than reproduce hidden benchmark tables.
References

Core Papers and Prior Episodes

Verified arXiv links are attached where available; earlier podcast episodes are linked directly.

Papers

Earlier Episodes