50× KV Cache Compression in Seconds

Fast KV Compaction via Attention Matching · Adam Zweiger, Xinghong Fu, Han Guo, Yoon Kim · MIT · February 2026
arXiv:2602.16284 KV Cache Attention Matching Latent Compaction

What Is the KV Cache?

Why memory grows punitively with context length during autoregressive inference

Per-Token Memory Accumulation

During inference, every token generates a Key and Value vector at every layer and every attention head. These are cached to avoid recomputation. The cache grows as: tokens × layers × heads × head_dim × 2 (K+V) × dtype_bytes

KV Cache Structure: Layers × Heads × Tokens Layers (32) Attention Heads (8 KV heads) Tokens (growing →) Cached K/V entries New token's entries (all layers, all heads)

Memory Scaling by Context Length

KV cache memory for a 32-layer, 8-KV-head model with head_dim=128 in bfloat16. Hover for exact values.

Context Length → Memory per Request
50×
Compression Ratio
~sec
Compaction Time
8–16 GB
Cache @ 64K tokens
160–320 MB
After AM Compaction

Why 50× Matters for Production

At 64K-token contexts, a single request consumes 8–16 GB of KV cache. Multiply by 16 concurrent users and your GPU memory budget is spent entirely on cache. At 50× compression, each request needs only 160–320 MB — enabling 50× more concurrent requests on the same hardware, or supporting million-token contexts within existing memory budgets.

The Compression Landscape

From token eviction to latent-space compaction: how the field evolved

Evolution of KV Cache Compression

2023 2024 2025 2026 Token Space Latent Space H2O SnapKV KV Merger CART- RIDGES AM paradigm shift →

Token-Space Methods

Select, evict, or merge original tokens. Fast but breaks at high compression.

Original (T tokens): After 50× eviction (2%): Most information discarded at 50×

H2O · SnapKV · KVMerger · PyramidKV

Latent-Space Compaction

Construct synthetic KV vectors that reproduce attention behavior. Preserves information at extreme ratios.

Original (T tokens): Compact cache (t synthetic vectors): Synthetic vectors encode full context behavior

Cartridges (hours) → AM (seconds)

Capability Matrix at 50× Compression

Hover cells for details

Attention Matching: The Algorithm

From gradient descent to closed-form solution — how AM bypasses iterative optimization

Compaction Pipeline

1Full KV Cache
→
2Reference Queries
→
3Gram Matrix
→
4Closed-Form Solve
→
5Compact Cache

Step 1: Full KV Cache

Run prefill on the full document (T tokens). Store keys K ∈ ℝ^(T×d) and values V ∈ ℝ^(T×d) at every layer and head. This is the starting point — the full, uncompressed representation.

Core Insight: Per-Head Decomposition

Cartridges: Holistic Optimization Full Model Forward + Backward Token Loss -log P(next) ∇ gradient descent (GPU-hours) Expensive: backprop through all layers AM: Per-Head Closed Form Head 1 Head 2 Head H ··· G = QᵀQ Gram matrix per head V* = G⁻¹QᵀV One matrix solve per head (seconds)

Gram Matrix: Reference Query Inner Products

The Gram matrix G = QᵀQ captures similarity between reference queries. Its inversion gives the optimal compact values. Diagonal dominance means reference queries are diverse — good for compaction quality.

Hover cells to see query similarity values

Speed: Gradient Descent vs Closed-Form

Cartridges GPU-hours Iterative gradient descent through full model AM (fast) ~seconds Single matrix solve per head AM (quality) ~1-2 min More reference queries, higher accuracy ~100× faster at comparable quality

Results & Pareto Analysis

Qwen3-4B on QuALITY benchmark, 50× compression, single H100 GPU

Pareto Frontier: Accuracy vs Compaction Time

Compaction Time (log scale) → Accuracy (%) → 30% 40% 50% 60% 70% 1s 10s 1min 10min 1hr+ Full context

QuALITY Accuracy at 50× Compression

Evaluation Gaps & Risk Register

Red = high risk · Orange = medium · Green = addressed

Honest Assessment

What's Real

The Pareto frontier over Cartridges at realistic compute budgets is genuine. The closed-form decomposition is mathematically sound. GPU-hours → seconds is a meaningful transition.

What's Missing

One model (4B params), one benchmark (retrieval-friendly), one compression ratio. No multi-hop reasoning, no code completion, no scaling beyond single H100. Reference query sensitivity uncharacterized.

References

Papers and methods discussed in this episode

  • Primary Fast KV Compaction via Attention Matching — Zweiger, Fu, Guo, Kim · MIT · 2026 · arXiv:2602.16284
  • Latent Cartridges: Compact Representations for Long Contexts — Eyuboglu et al. · Stanford, CMU · 2025 · Episode 37
  • Eviction H2O: Heavy-Hitter Oracle for Efficient Generative Inference — Zhang et al. · 2023
  • Selection SnapKV: LLM Knows What You Are Looking For Before Generation — Li et al. · 2024
  • Merging KVMerger: Merging Key-Value Pairs for Efficient LLM Inference — Wang et al. · 2024
  • Queries KVzip: Query-Dependent KV Cache Compression — Kim et al. · 2025
  • Prefix Prefix-Tuning: Optimizing Continuous Prompts — Li & Liang · Stanford · 2021
  • Attn FlashAttention: Fast and Memory-Efficient Attention — Dao et al. · Stanford · 2022
  • Quant GPTQ: Accurate Post-Training Quantization — Frantar et al. · IST Austria · 2022
  • Sparse SparseGPT: Massive Language Models Can Be Pruned in One Shot — Frantar & Alistarh · IST Austria · 2023
  • Infra PagedAttention / vLLM — Kwon et al. · UC Berkeley · 2023
  • SVD Thin Keys Full Values: Compressing KV Cache via SVD
  • Distill KV-Distill: Nearly Lossless Learnable Context Compression
  • Layer PyramidKV: Dynamic KV Cache Compression Based on Pyramidal Information Funneling
  • Bench QuALITY: Question Answering with Long Input Texts — Pang et al. · NYU · 2022
  • Model Qwen3 Model Family — Alibaba · 2025

Episode Context

This visualization accompanies Episode 45 of AI Post Transformers. Related episodes: Ep. 37 (Cartridges), long-context dichotomy discussion.