Serving Path
Switch the deployment mode to see how interference collapses back into a shared pool when prefill and decode fight for the same accelerators.
Mooncake reframes long-context LLM serving around a moving cache, not just busy GPUs. These views track why TTFT and TBT pull the system in different directions, how tiered KV placement changes routing, and where Mooncake’s scheduler wins or refuses work.
Mooncake separates the work that wins the first token from the work that keeps later tokens smooth. The split only pays if the transfer tax stays smaller than the prefill you avoided.
Switch the deployment mode to see how interference collapses back into a shared pool when prefill and decode fight for the same accelerators.
Both rows render the same request class. The active focus amplifies either the first-token path or the steady decode rhythm.
Mooncake treats cache blocks like movable assets. Reuse length, access tier, and fetch cost decide whether a request should hydrate, recompute, migrate, or replicate.
Hover blocks to inspect where reuse sits. Blue cells are cold. Orange and red cells are the blocks the scheduler wants to keep nearest to decode.
The arrows show where bytes move for the selected scenario. Wider links mean more of the request state takes that path before decode can start.
Mooncake does not simply pick the emptiest GPU. It scores candidate prefill/decode pairs by reusable prefix length, remote transfer cost, queue delay, and whether decode will still hit token latency after admission.
Step through the scheduler path. The winning route changes when the shared prefix is too short to justify remote fetch, or when decode headroom is already gone.
Reactive admission waits until decode is visibly full. Predictive admission closes earlier, which avoids the queue oscillation Mooncake argues against.
The paper’s strongest evidence is not “faster everywhere.” It is “much better when the workload is long, prefix-heavy, and overloaded enough that cache placement and admission actually matter.”
Public datasets move modestly. Synthetic long-context reuse can explode upward. Real trace replay lands between those poles and shows the architecture surviving deeper overload.
The key curve is successful requests that still meet latency objectives. Mooncake’s line stays higher for longer because it routes by cached state and rejects earlier when decode would collapse.