AI Post Transformers · Episode Companion

Global Memory Bloat in Long-Context LLM Serving

📄 arXiv:2607.02574 Texas Tech University Jie Li, Tongyang Wang, Yong Chen · 2026 30+ systems surveyed
40 GB
KV cache, 1 req · 128K ctx · 70B
~1.5 PB
cumulative reads, 32K-tok decode
35
systems classified
2 / 5
lifetime categories with zero systems

Prefill vs. Decode: two phases, opposite bottlenecks

Prefill writes the KV cache in one compute-bound pass over the prompt. Decode rereads the entire accumulated cache at every single generated token — that's the bandwidth trap.

Serving pipeline

Hover the boxes. Prefill is throughput-bound; decode is bandwidth- and latency-bound, since it rereads the whole cache every step.

KV footprint by attention variant — 70B model, FP16

Grouped-query attention (8 KV heads) cuts footprint vs. full multi-head attention proportionally. DeepSeek-V2/V3's Multi-head Latent Attention (MLA) compresses further into one low-rank latent — reported 5–13× smaller than GQA at matched quality.

MHA (full heads) GQA (8 KV heads, industry-standard) MLA (DeepSeek compressed latent)

Capacity is a snapshot. Bandwidth is a cumulative bill.

The 40 GB number is resident memory at one instant, scaling linearly with sequence length. But decode rereads the growing cache at every step — cumulative traffic scales closer to quadratically.

The formula, factor by factor

KV bytes = 2 (K & V) × layers × KV heads × head dim × sequence length × bytes/elem. Hover a step to see its multiplier land on the running total.

Capacity vs. bandwidth — same growth, different physics

Toggle between resident footprint (linear in context length) and cumulative bytes moved across a full decode run (superlinear in tokens generated).

Four axes, thirty-five systems — and two empty cells

Locality × lifetime, mock counts illustrating the paper's classification. B1 (per-session) and B4 (durable-recoverable) contain zero real-world systems out of thirty-five surveyed — an empirical hole, not a modeling choice.

Locality × Lifetime — system counts

Hover a cell for detail. Dashed orange cells are the empty B1/B4 columns.

few systems moderate many systems 0 systems — design gap

Axis legend

Locality: A0 local-only · A1 direct remote · A2 shared store · A3 shared memory (CXL).
Lifetime: B0 per-request · B1 per-session (empty) · B2 cross-session · B3 opportunistic retention (modifier) · B4 durable-recoverable (empty).
Ownership: C0 per-worker → C3 shared-memory convention. Transport: D-local, NVLink, RDMA, transfer libs, CXL, host DRAM, SSD, SmartNICs.

Thirty-five systems collapse into five archetypes

Scoring every system on the four-way tuple, instead of scattering across the full space, they clump into five repeated shapes — unevenly populated and unevenly reviewed.

Systems per archetype

Toggle total count vs. peer-reviewed subset. Memory-pool's three systems are zero-for-three on peer review.

References