Still compresses a transformer's key-value cache into a fixed-size representation in a single forward pass — manufacturing new compact keys and values instead of just picking survivors, with a careful RoPE workaround so blended positions don't destabilize.
Compaction splits along what you keep (selection vs. synthesis) and when the work happens (per-context vs. amortized). Selection has been amortized before; synthesis had not — until Still. Hover a cell for detail.
Cache size grows linearly with context. Past a point it exceeds the model's own weights on the GPU — the exact regime multi-day agents and repo-scale reasoning now push into.
Single-pass: prefill once, compact once. Iterative: a scheduled re-compaction that stays cheap because each step is just one more forward pass, never a training loop.
A small bank of learned latents (Perceiver-style) distills many cached tokens into a fixed-size output — many-to-few, not a subset.
Cached keys are already rotated by position. Still un-rotates into a position-free frame, blends safely, then re-rotates the manufactured output at new, freely-assigned positions.
~120k extractive multiple-choice items across four domains, ~1B tokens at 8k context, filtered so the question is unanswerable without the context.
Teacher reads the full cache, student reads the compact one. Forward KL divergence, masked to answer tokens. The base model stays frozen throughout — gradients flow only into the Perceiver compactor.
Still is the only method that stays fast without trading away accuracy as the compression ratio climbs from 8x to 200x.
Still beats KV-Distill by 8–22 accuracy points across most of the grid, widening at longer context.
The 8k-trained checkpoint collapses to 1.5% at 128k — below the no-context floor. A 32k-trained checkpoint degrades far more gracefully. Training horizon, not capacity or budget, is the binding constraint.
On multi_lexsum, Still recovers 74–95% of the full-context gain through 64k, still 59% at 128k — ahead of Attention Matching and KV-Distill at every length.
Capacity-bound serving (many concurrent long-horizon agent sessions): Still helps today — tens of MiB instead of tens of GiB per session means far more sessions resident in the same GPU memory.
~144 KiB/token summed across layers and heads means one 128k-token session can eat tens of GiB. Still's compact cache is tens of MiB regardless of original length — a capacity story.
The one real serving measurement: the compact path is slower to first token than the full cache. The abstract's "favorable side of the speed-quality frontier" measures offline compaction time, not this.
The strongest quality reference, Cartridges, never gets a head-to-head number. The generation-aligned checkpoints get 50 extra fine-tuning steps per target task before the HELMET/LongBench numbers are measured — a fine-tuned descendant, not quite the same claim as "the compactor generalizes." Iterative compaction collapses past its trained horizon, the ratio is fixed rather than a real budget, needle retrieval is near zero, and there's no agentic tool-use evaluation despite that being the paper's own opening scenario.