Why a KV Cache Exists
Without caching, every new token re-attends to the entire history — quadratic cost per step. The KV cache stores keys/values from prefill so decode only pays linear cost.
Three Paradigms, One Problem
Same goal — fit the cache in memory — three incompatible bets on what to sacrifice.
Time-to-First-Token: Where Prefill Falls Over
GPU memory required to complete prefill, as a function of prompt length (log scale). vLLM scales cleanly to the full 128K window. H2O and InfiniGen both blow past the 80GB H100 budget around 10K tokens.
FlashAttention-2 Rescue — But Only for One of Them
FA-2 fuses attention into SRAM tiles so the full score matrix never materializes. That helps InfiniGen (which just needs attention output) but does nothing for H2O (which needs the materialized softmax scores to rank tokens for eviction).
Resource Footprint (Llama-3.1-8B, batch 16→96)
vLLM holds ~72GB flat regardless of batch size. H2O stays under 40GB even at batch 96 — roughly half of vLLM.
Throughput by Batch Size (tok/s)
Time to Generate 8,192 Tokens
InfiniGen serializes CPU↔GPU transfer on every decode step — it compounds badly at long generations.
Accuracy Degradation vs. Baseline vLLM
At 0.3 — the operating point picked for the whole study — both land within 10% on aggregate benchmarks. Drop to 0.1 and H2O falls off a cliff.
Fact-Retention Task (recall %, inject-then-ask-later)
Same 0.3 budget as the "equivalent" accuracy claim above — yet H2O collapses on long-range retention while InfiniGen tracks close to baseline. Hover a cell for the exact figure.
The Structural Gap: GQA Shrinks the Cache, Not the Score Matrix
GQA cuts steady-state KV cache size by sharing heads. But eviction-based methods need the transient attention-score matrix (Q·Kᵀ) materialized during prefill to decide what to evict — and GQA doesn't touch that at all.