A KV cache is fastest to read on the GPU and slowest on disk — but disk is where the capacity, and most cache hits, actually live. Toggle to compare bandwidth vs. capacity.
Hover a slice — most reused KV cache isn't sitting in fast GPU memory, it's on the tier with the narrowest pipe.
Every query token attends to every key/value token before it. Row i does i+1 units of attention work — hover a cell.
This asymmetry is the whole paper: a KV tensor costs the same to move regardless of position, but computing it fresh gets more expensive the further into the sequence you go. Hover a point.
compute_ptr walks forward from chunk 0 on the GPU. io_ptr
walks backward from the last chunk over the I/O path. The prefill finishes the
instant they cross — two tunnel crews digging from opposite ends of the sequence.
Cake's own prefill doesn't get GPU priority — other decode and prefill traffic go first. Toggle to see how the I/O thread absorbs the slack instead of stalling.
Same chart, five views. Cake's advantage grows with more idle compute, smaller caches (GQA, quantization), and longer sequences.
Every table and figure in Section 5 reports TTFT. Nothing else — no accuracy, no perplexity, no test of whether a prefix computed fresh and a suffix loaded from disk (or quantized) produce a coherent output.
MLA (DeepSeek-V2) compresses KV cache far past GQA. Less to move over I/O means less for the bidirectional scheduler to parallelize against — Cake was never tested on it.
Precision boundary, untrusted: Section 5.5 stitches a full-precision computed prefix to an 8-bit or 3-bit quantized loaded suffix, mid-sequence, with no check. CacheBlend (Yao, Li, Liu et al., U. Chicago, 2024) exists specifically to correct this kind of chunk-fusion deviation — Cake cites it, but not for this.
Simulated I/O, not real contention: bandwidth is a fixed artificial delay, not a measured link. Mooncake (Qin et al., Moonshot AI/Kimi, 2024) moves KV cache across real network links at production scale — the regime where a computed merge point might not hold.
Single-node, single-request only: the calculated meet-in-the-middle point is never tested under multi-tenant or cross-node contention.
Failure mode: short sequences with lopsided resources (e.g. two A100s vs. a slow 7 Gb/s link) — Cake still ships a chunk over I/O out of habit, turning into pure overhead. The authors sketch a single-resource fallback but don't build it.