arXiv:2410.03065

Compute or Load? Cake's Smart KV Cache Scheduler

Compute Or Load KV Cache? Why Not Both? — Shuowei Jin, Xueshen Liu, Qingzhao Zhang, Z. Morley Mao · University of Michigan · posted arXiv Oct 2024, revised Feb 2026. Introduces Cake, a two-pointer scheduler that computes and loads a prompt's KV cache simultaneously from opposite ends.

30s
prefill time for a 72,000-token input, Llama2-70B on one A100 — dead air before token one
~80%
of prefix-cache hits land on the slowest, disk-bound storage tier (AttentionStore, USENIX ATC 2024)
50%+
cost reduction reported by OpenAI, Anthropic & DeepSeek on prefix-cache hits in production

Memory hierarchy: fast is small, big is slow

A KV cache is fastest to read on the GPU and slowest on disk — but disk is where the capacity, and most cache hits, actually live. Toggle to compare bandwidth vs. capacity.

Where cache hits actually land

Hover a slice — most reused KV cache isn't sitting in fast GPU memory, it's on the tier with the narrowest pipe.

Why compute cost rises: the causal mask

Every query token attends to every key/value token before it. Row i does i+1 units of attention work — hover a cell.

Hover any cell in the grid.

Compute rises, I/O stays flat

This asymmetry is the whole paper: a KV tensor costs the same to move regardless of position, but computing it fresh gets more expensive the further into the sequence you go. Hover a point.

Hover a compute (green) or I/O (purple) point.
compute cost / chunk I/O cost / chunk

Meet-in-the-middle: two pointers, two threads

compute_ptr walks forward from chunk 0 on the GPU. io_ptr walks backward from the last chunk over the I/O path. The prefill finishes the instant they cross — two tunnel crews digging from opposite ends of the sequence.

computed on GPU loaded via I/O pending

Adaptive priority under real GPU contention

Cake's own prefill doesn't get GPU priority — other decode and prefill traffic go first. Toggle to see how the I/O thread absorbs the slack instead of stalling.

TTFT speedup, sliced five ways

Same chart, five views. Cake's advantage grows with more idle compute, smaller caches (GQA, quantization), and longer sequences.

What this paper actually measured gap

Every table and figure in Section 5 reports TTFT. Nothing else — no accuracy, no perplexity, no test of whether a prefix computed fresh and a suffix loaded from disk (or quantized) produce a coherent output.

Smaller caches shrink Cake's whole reason to exist

MLA (DeepSeek-V2) compresses KV cache far past GQA. Less to move over I/O means less for the bidirectional scheduler to parallelize against — Cake was never tested on it.

Open questions worth tracking

Precision boundary, untrusted: Section 5.5 stitches a full-precision computed prefix to an 8-bit or 3-bit quantized loaded suffix, mid-sequence, with no check. CacheBlend (Yao, Li, Liu et al., U. Chicago, 2024) exists specifically to correct this kind of chunk-fusion deviation — Cake cites it, but not for this.

Simulated I/O, not real contention: bandwidth is a fixed artificial delay, not a measured link. Mooncake (Qin et al., Moonshot AI/Kimi, 2024) moves KV cache across real network links at production scale — the regime where a computed merge point might not hold.

Single-node, single-request only: the calculated meet-in-the-middle point is never tested under multi-tenant or cross-node contention.

Failure mode: short sequences with lopsided resources (e.g. two A100s vs. a slow 7 Gb/s link) — Cake still ships a chunk over I/O out of habit, turning into pure overhead. The authors sketch a single-resource fallback but don't build it.

References

  1. Compute Or Load KV Cache? Why Not Both? — Shuowei Jin, Xueshen Liu, Qingzhao Zhang, Z. Morley Mao, 2024. arxiv.org/abs/2410.03065
  2. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Qin, R. et al. (Moonshot AI / Kimi), 2024. Google Scholar
  3. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI, 2024. Google Scholar
  4. CacheBlend: Fast Large Language Model Serving with Cached Knowledge Fusion — Yao, J., Li, H., Liu, Y., et al., 2024. Google Scholar
  5. Preble: Efficient Distributed Prompt Scheduling for LLM Serving — Srivatsa, V. et al., 2024. Google Scholar