Fast KV Compaction via Attention Matching · Adam Zweiger, Xinghong Fu, Han Guo, Yoon Kim · MIT · February 2026 arXiv:2602.16284KV CacheAttention MatchingLatent Compaction
What Is the KV Cache?
Why memory grows punitively with context length during autoregressive inference
Per-Token Memory Accumulation
During inference, every token generates a Key and Value vector at every layer and every attention head. These are cached to avoid recomputation. The cache grows as: tokens × layers × heads × head_dim × 2 (K+V) × dtype_bytes
Memory Scaling by Context Length
KV cache memory for a 32-layer, 8-KV-head model with head_dim=128 in bfloat16. Hover for exact values.
50×
Compression Ratio
~sec
Compaction Time
8–16 GB
Cache @ 64K tokens
160–320 MB
After AM Compaction
Why 50× Matters for Production
At 64K-token contexts, a single request consumes 8–16 GB of KV cache. Multiply by 16 concurrent users and your GPU memory budget is spent entirely on cache. At 50× compression, each request needs only 160–320 MB — enabling 50× more concurrent requests on the same hardware, or supporting million-token contexts within existing memory budgets.
The Compression Landscape
From token eviction to latent-space compaction: how the field evolved
Evolution of KV Cache Compression
Token-Space Methods
Select, evict, or merge original tokens. Fast but breaks at high compression.
H2O · SnapKV · KVMerger · PyramidKV
Latent-Space Compaction
Construct synthetic KV vectors that reproduce attention behavior. Preserves information at extreme ratios.
Cartridges (hours) → AM (seconds)
Capability Matrix at 50× Compression
Attention Matching: The Algorithm
From gradient descent to closed-form solution — how AM bypasses iterative optimization
Compaction Pipeline
1Full KV Cache
→
2Reference Queries
→
3Gram Matrix
→
4Closed-Form Solve
→
5Compact Cache
Step 1: Full KV Cache
Run prefill on the full document (T tokens). Store keys K ∈ ℝ^(T×d) and values V ∈ ℝ^(T×d) at every layer and head. This is the starting point — the full, uncompressed representation.
Core Insight: Per-Head Decomposition
Gram Matrix: Reference Query Inner Products
The Gram matrix G = QᵀQ captures similarity between reference queries. Its inversion gives the optimal compact values. Diagonal dominance means reference queries are diverse — good for compaction quality.
Speed: Gradient Descent vs Closed-Form
Results & Pareto Analysis
Qwen3-4B on QuALITY benchmark, 50× compression, single H100 GPU
Pareto Frontier: Accuracy vs Compaction Time
QuALITY Accuracy at 50× Compression
Evaluation Gaps & Risk Register
Honest Assessment
What's Real
The Pareto frontier over Cartridges at realistic compute budgets is genuine. The closed-form decomposition is mathematically sound. GPU-hours → seconds is a meaningful transition.
What's Missing
One model (4B params), one benchmark (retrieval-friendly), one compression ratio. No multi-hop reasoning, no code completion, no scaling beyond single H100. Reference query sensitivity uncharacterized.
References
Papers and methods discussed in this episode
PrimaryFast KV Compaction via Attention Matching — Zweiger, Fu, Guo, Kim · MIT · 2026 · arXiv:2602.16284
LatentCartridges: Compact Representations for Long Contexts — Eyuboglu et al. · Stanford, CMU · 2025 · Episode 37
EvictionH2O: Heavy-Hitter Oracle for Efficient Generative Inference — Zhang et al. · 2023
SelectionSnapKV: LLM Knows What You Are Looking For Before Generation — Li et al. · 2024
MergingKVMerger: Merging Key-Value Pairs for Efficient LLM Inference — Wang et al. · 2024
QueriesKVzip: Query-Dependent KV Cache Compression — Kim et al. · 2025