arXiv:2606.07878

KV Cache Compaction Beyond Selection

Introducing Still

Charles O'Neill et al. · Baseten · posted June 5, 2026

Still compresses a transformer's key-value cache into a fixed-size representation in a single forward pass — manufacturing new compact keys and values instead of just picking survivors, with a careful RoPE workaround so blended positions don't destabilize.

8x–200x compression sweep 8k–128k context range Qwen3 4B–32B + Gemma-3 Perceiver-based compactor

Two-axis taxonomy of KV compaction

Compaction splits along what you keep (selection vs. synthesis) and when the work happens (per-context vs. amortized). Selection has been amortized before; synthesis had not — until Still. Hover a cell for detail.

Why the cache becomes the bottleneck

Cache size grows linearly with context. Past a point it exceeds the model's own weights on the GPU — the exact regime multi-day agents and repo-scale reasoning now push into.

Per-layer compaction pipeline

Single-pass: prefill once, compact once. Iterative: a scheduled re-compaction that stays cheap because each step is just one more forward pass, never a training loop.

Latent queries cross-attend into the cache

A small bank of learned latents (Perceiver-style) distills many cached tokens into a fixed-size output — many-to-few, not a subset.

The RoPE workaround

Cached keys are already rotated by position. Still un-rotates into a position-free frame, blends safely, then re-rotates the manufactured output at new, freely-assigned positions.

Synthetic training corpus

~120k extractive multiple-choice items across four domains, ~1B tokens at 8k context, filtered so the question is unanswerable without the context.

Only the compactor learns

Teacher reads the full cache, student reads the compact one. Forward KL divergence, masked to answer tokens. The base model stays frozen throughout — gradients flow only into the Perceiver compactor.

Accuracy across the compression sweep

Still is the only method that stays fast without trading away accuracy as the compression ratio climbs from 8x to 200x.

RULER vs. KV-Distill (matched training)

Still beats KV-Distill by 8–22 accuracy points across most of the grid, widening at longer context.

Iterative compaction: collapse past the training horizon

The 8k-trained checkpoint collapses to 1.5% at 128k — below the no-context floor. A 32k-trained checkpoint degrades far more gracefully. Training horizon, not capacity or budget, is the binding constraint.

HELMET summarization: % of full-context gain recovered

On multi_lexsum, Still recovers 74–95% of the full-context gain through 64k, still 59% at 128k — ahead of Attention Matching and KV-Distill at every length.

Capacity-bound serving (many concurrent long-horizon agent sessions): Still helps today — tens of MiB instead of tens of GiB per session means far more sessions resident in the same GPU memory.

Memory footprint per session (log scale)

~144 KiB/token summed across layers and heads means one 128k-token session can eat tens of GiB. Still's compact cache is tens of MiB regardless of original length — a capacity story.

vLLM time-to-first-token (Appendix B.1, single H200)

The one real serving measurement: the compact path is slower to first token than the full cache. The abstract's "favorable side of the speed-quality frontier" measures offline compaction time, not this.

What's still open

The strongest quality reference, Cartridges, never gets a head-to-head number. The generation-aligned checkpoints get 50 extra fine-tuning steps per target task before the HELMET/LongBench numbers are measured — a fine-tuned descendant, not quite the same claim as "the compactor generalizes." Iterative compaction collapses past its trained horizon, the ratio is fixed rather than a real budget, needle retrieval is near zero, and there's no agentic tool-use evaluation despite that being the paper's own opening scenario.

References

  1. O'Neill et al. — KV Cache Compaction Beyond Selection: Introducing Still (2026) — arXiv:2606.07878
  2. Zhang, Sheng, Zhou et al. — H2O: Heavy-Hitter Oracle for Efficient Generative Inference (2023) — Scholar
  3. Xiao, Tian, Chen et al. — Efficient Streaming Language Models with Attention Sinks (2023) — Scholar
  4. Mu, Li, Goodman — Learning to Compress Prompts with Gist Tokens (2023) — Scholar
  5. Eyuboglu, Ehrlich, Arora, Guha et al. — Cartridges: Long Context via Self-Study (2025) — Scholar
  6. Lee, Lee, Kim, Kosiorek et al. — Set Transformer (2019) — Scholar
  7. Jaegle, Gimeno, Brock et al. — Perceiver: General Perception with Iterative Attention (2021) — Scholar
  8. Alayrac, Donahue, Luc, Miech et al. — Flamingo: a Visual Language Model for Few-Shot Learning (2022) — Scholar
  9. Li, Li, Savarese, Hoi — BLIP-2 (2023) — Scholar
  10. Diaz — Learned structure in cartridges: Keys as shareable routers (2025) — Scholar
  11. Zweiger, Fu, Guo, Kim — Fast KV compaction via attention matching (2026) — Scholar
  12. Chari, Qin, Van Durme — KV-Distill (2025) — Scholar
  13. Moschella, Manduchi, Sener — Learning to evict from key-value cache (KVP) (2026) — Scholar
  14. DeepSeek-AI — DeepSeek-V4 (2026) — Scholar