Prefill writes the KV cache in one compute-bound pass over the prompt. Decode rereads the entire accumulated cache at every single generated token — that's the bandwidth trap.
Hover the boxes. Prefill is throughput-bound; decode is bandwidth- and latency-bound, since it rereads the whole cache every step.
Grouped-query attention (8 KV heads) cuts footprint vs. full multi-head attention proportionally. DeepSeek-V2/V3's Multi-head Latent Attention (MLA) compresses further into one low-rank latent — reported 5–13× smaller than GQA at matched quality.
The 40 GB number is resident memory at one instant, scaling linearly with sequence length. But decode rereads the growing cache at every step — cumulative traffic scales closer to quadratically.
KV bytes = 2 (K & V) × layers × KV heads × head dim × sequence length × bytes/elem. Hover a step to see its multiplier land on the running total.
Toggle between resident footprint (linear in context length) and cumulative bytes moved across a full decode run (superlinear in tokens generated).
Locality × lifetime, mock counts illustrating the paper's classification. B1 (per-session) and B4 (durable-recoverable) contain zero real-world systems out of thirty-five surveyed — an empirical hole, not a modeling choice.
Hover a cell for detail. Dashed orange cells are the empty B1/B4 columns.
Locality: A0 local-only · A1 direct remote · A2 shared store · A3 shared memory (CXL).
Lifetime: B0 per-request · B1 per-session (empty) · B2 cross-session · B3 opportunistic retention (modifier) · B4 durable-recoverable (empty).
Ownership: C0 per-worker → C3 shared-memory convention. Transport: D-local, NVLink, RDMA, transfer libs, CXL, host DRAM, SSD, SmartNICs.
Scoring every system on the four-way tuple, instead of scattering across the full space, they clump into five repeated shapes — unevenly populated and unevenly reviewed.
Toggle total count vs. peer-reviewed subset. Memory-pool's three systems are zero-for-three on peer review.