KV Cache Management Faces Its First Head-to-Head Test

arXiv:2604.05012 4× H100 Mamo, Kogiou, Yi, Yu · Florida State University · 2026

The first side-by-side test of vLLM's PagedAttention, H2O's static sparsification, and InfiniGen's dynamic CPU-offload — on identical hardware. Headline finding: H2O and InfiniGen both hit out-of-memory errors around 10,000 tokens, under 10% of a claimed 128K context window.

Why a KV Cache Exists

Without caching, every new token re-attends to the entire history — quadratic cost per step. The KV cache stores keys/values from prefill so decode only pays linear cost.

Three Paradigms, One Problem

Same goal — fit the cache in memory — three incompatible bets on what to sacrifice.

4×
GQA cache reduction (32→8 KV heads, Llama-3.1)
30×+
context growth, 4K → 128K, since GQA landed
<10%
of 128K window survived by H2O / InfiniGen baseline

Time-to-First-Token: Where Prefill Falls Over

GPU memory required to complete prefill, as a function of prompt length (log scale). vLLM scales cleanly to the full 128K window. H2O and InfiniGen both blow past the 80GB H100 budget around 10K tokens.

FlashAttention-2 Rescue — But Only for One of Them

FA-2 fuses attention into SRAM tiles so the full score matrix never materializes. That helps InfiniGen (which just needs attention output) but does nothing for H2O (which needs the materialized softmax scores to rank tokens for eviction).

Resource Footprint (Llama-3.1-8B, batch 16→96)

vLLM holds ~72GB flat regardless of batch size. H2O stays under 40GB even at batch 96 — roughly half of vLLM.

Throughput by Batch Size (tok/s)

Time to Generate 8,192 Tokens

InfiniGen serializes CPU↔GPU transfer on every decode step — it compounds badly at long generations.

Accuracy Degradation vs. Baseline vLLM

At 0.3 — the operating point picked for the whole study — both land within 10% on aggregate benchmarks. Drop to 0.1 and H2O falls off a cliff.

Fact-Retention Task (recall %, inject-then-ask-later)

Same 0.3 budget as the "equivalent" accuracy claim above — yet H2O collapses on long-range retention while InfiniGen tracks close to baseline. Hover a cell for the exact figure.

The Structural Gap: GQA Shrinks the Cache, Not the Score Matrix

GQA cuts steady-state KV cache size by sharing heads. But eviction-based methods need the transient attention-score matrix (Q·Kᵀ) materialized during prefill to decide what to evict — and GQA doesn't touch that at all.