1 00:00:01,000 --> 00:00:25,425 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper is 'Still: Amortized KV Cache Compaction in a Single Forward Pass' — Charles O'Neill et al., five authors total, out of Baseten, posted to arXiv on June 5th, 2026. Ada, sell this to me — cache management doesn't sound thrilling on paper. 2 00:00:25,425 --> 00:00:44,799 [Dr. Ada Shannon] Honestly, it's sexier than it sounds. What pulled me in wasn't a benchmark chart, it's how carefully they map the design space before proposing anything — what's tried, what's structurally impossible, where nobody's built anything. Then they walk straight into that gap. That front-loaded rigor is rare — once they land on the method, you already understand why it should work. 3 00:00:44,799 --> 00:00:51,750 [Hal Turing] Start us at the bottom, then — what actually is the KV cache, and why's it such a bottleneck now? 4 00:00:51,750 --> 00:01:13,674 [Dr. Ada Shannon] Every transformer keeps a running memory of every token it's already seen — a Key and Value vector per token, per layer. That's the cache, so generating token ten thousand doesn't mean recomputing attention over the prior nine thousand ninety-nine, you just read them back out. It grows linearly with context length, though — and with multi-day agents and repo-scale reasoning now, that cache can balloon past the size of the model's own weights, the real bottleneck on the GPU. 5 00:01:13,674 --> 00:01:20,300 [Hal Turing] So what's the playbook today, before this paper — what do people reach for when the cache gets too big? 6 00:01:20,300 --> 00:01:35,800 [Dr. Ada Shannon] Right now you've got two bad options. Keep the full cache and eat the linear memory growth till it doesn't fit. Or abandon the model's own internal state — fine-tune it in, bolt on retrieval-augmented generation, or hand it a summary instead of— 7 00:01:35,800 --> 00:01:49,250 [Hal Turing] Oh wait wait wait — so it's really that binary right now? Either eat the full linear memory bill, or throw away the model's own internal representations and hope RAG or a summary caught what mattered? 8 00:01:49,250 --> 00:02:25,075 [Dr. Ada Shannon] Pretty much — that's the gap this paper's aimed at. Compacting the cache instead of abandoning it splits along two axes. First: how do you shrink it? Selection methods pick a subset of original entries and drop the rest — always literal survivors. Synthesis methods manufacture brand-new key and value vectors blending many positions, so what's left isn't a subset of anything. Second: when does the work happen? Per-context methods optimize fresh per document at inference — expressive, slow. Amortized methods train a module once, offline, so inference is one forward pass. Trained-in bakes compression into pretraining itself. 9 00:02:25,075 --> 00:02:31,400 [Hal Turing] And 'amortized' here means what, exactly — some tiny trained network sitting in between? 10 00:02:31,400 --> 00:03:03,925 [Dr. Ada Shannon] That's basically it, but the shape matters. Still's module is a Perceiver — a small bank of learned query vectors that cross-attend into a huge input and distill it to a fixed-size output regardless of length. That's from Jaegle and colleagues' 2021 Perceiver paper, later reused in DeepMind's Flamingo as the Perceiver Resampler, compressing image features into a handful of tokens. Same instinct behind variational autoencoders and sparse autoencoders — replace solving a fresh optimization per example with training one encoder that does it once. 11 00:03:03,925 --> 00:03:11,450 [Hal Turing] And position must matter a ton here too — you can't just mash vectors from different spots together blindly. 12 00:03:11,450 --> 00:03:29,950 [Dr. Ada Shannon] Right, that's RoPE — Rotary Position Embedding, from Su and colleagues' RoFormer paper. It encodes position by rotating a token's key vector, rather than adding a separate position signal. Elegant, but it means keys from different positions aren't directly combinable without accounting for that rotation. 13 00:03:29,950 --> 00:03:39,475 [Hal Turing] But hold on — amortized selection already exists, a learned scorer picking tokens in one pass. Doesn't that already fix the slow part? 14 00:03:39,475 --> 00:03:54,775 [Dr. Ada Shannon] I actually disagree, Hal. Amortizing when you decide doesn't change what you're allowed to keep. Selection is always bound to rows that already existed — however fast you pick, you can't blend two facts on different tokens into one slot. 15 00:03:54,775 --> 00:04:00,300 [Hal Turing] Does that ceiling actually bite in practice, or is it mostly theoretical? 16 00:04:00,300 --> 00:04:18,700 [Dr. Ada Shannon] It bites exactly where it matters — dense information, tight budgets. Nobody had married 'manufacture new content' with 'one forward pass' until this paper. Selection got amortized; synthesis stayed slow and per-context. Still is aimed straight at that empty quadrant, and that's where we're headed. 17 00:04:18,700 --> 00:04:26,125 [Hal Turing] Okay, so how does that actually play out per layer, Ada — walk me through what's actually happening to the cache. 18 00:04:26,125 --> 00:04:47,300 [Dr. Ada Shannon] Each layer runs its own version: the latent queries cross-attend into that layer's full key-value cache, refine through self-attention, then project out to compact keys and values — all in one forward pass, no gradient steps, no reference queries. Same move as Flamingo's Perceiver Resampler, Alayrac et al., DeepMind, 2022, just per-layer, on the cache itself, not once on a vision encoder. 19 00:04:47,300 --> 00:04:55,126 [Hal Turing] And there's a wrinkle with position too, right? Earlier you mentioned RoPE gets stripped out before any of this happens. 20 00:04:55,126 --> 00:05:16,801 [Dr. Ada Shannon] Right — cached keys are already RoPE-rotated, so identical content looks different by position. Blend content from a dozen positions using already-rotated keys and you're mixing content with phase — unstable. So Still un-rotates keys into a position-free frame, runs its own internal RoPE on the latents at evenly spaced positions, then re-rotates the output keys at whatever position you write them into the new cache. 21 00:05:16,801 --> 00:05:27,001 [Hal Turing] Wait, hold on — so the compact cache doesn't even inherit the original ordering? You're free to relabel positions however you like after compaction? 22 00:05:27,001 --> 00:05:48,926 [Dr. Ada Shannon] Exactly — and that's what makes iterative use possible. Single-pass is prefill once, compact once. For long-horizon agents they run it on a schedule instead: a chunk arrives, prefilled against the compacted history plus one raw lookahead chunk left deliberately uncompacted, so new tokens aren't squashed before the model's attended to them. Compaction being a forward pass is what makes calling it repeatedly, mid-trajectory, cheap. 23 00:05:48,926 --> 00:05:55,701 [Hal Turing] So how do you even train something like this without a mountain of labeled long-context data? 24 00:05:55,701 --> 00:06:24,426 [Dr. Ada Shannon] Synthetic, and clever about it. Four domains — financial filings, Gutenberg literature, legal documents, code — extractive multiple-choice questions from random sub-chunks, verified by the same frozen base model, filtering out anything answerable without context. About 120,000 items, a billion tokens at 8k. Loss is forward KL divergence — teacher reads the full cache, student reads the compact one, masked to answer tokens only. The compactor's the only thing that updates; base model stays frozen throughout. 25 00:06:24,426 --> 00:06:28,401 [Hal Turing] Okay, so does it actually deliver? What's the headline? 26 00:06:28,401 --> 00:07:02,726 [Dr. Ada Shannon] Across an 8x-to-200x compression sweep and 8k-to-128k contexts, Still's the only method that stays fast without giving up accuracy — everything else trades one for the other. Same recipe transfers untouched across the Qwen3 dense family, 4B to 32B, the 30B-A3B mixture-of-experts, and Gemma-3's mixed sliding-window/global stack, where they just compact the global layers. And on RULER, matched-training setup, it beats KV-Distill — Chari, Qin, and Van Durme, Johns Hopkins, 2025 — by 8 to 22 accuracy points across most of the grid. 27 00:07:02,726 --> 00:07:17,752 [Hal Turing] And the iterative results are honestly the wildest part to me — a compactor trained at 16k soaking up 128k of trajectory, way past its physical latent capacity. That's real extrapolation. 28 00:07:17,752 --> 00:07:36,978 [Dr. Ada Shannon] I'd push back on calling that clean extrapolation, Hal. The 8k-trained checkpoint, pushed to 128k, collapses to 1.5 percent — below the no-context floor. That's not graceful degradation, it's the cache actively getting corrupted once you stack more compactions than it ever saw in training. 29 00:07:36,978 --> 00:07:46,978 [Hal Turing] Sure, but that's an unfair test, isn't it — nobody said an 8k-trained model should generalize sixteen times past its horizon for free. 30 00:07:46,978 --> 00:08:06,503 [Dr. Ada Shannon] No no — that's exactly the finding worth sitting with. The 8k checkpoint actually retains a larger raw cache at 128k than the 32k-trained one, and still collapses. Not a capacity or budget problem — training horizon is the binding constraint, full stop, and that's a real limitation. 31 00:08:06,503 --> 00:08:18,403 [Hal Turing] Fair — common ground: impressive within range, brittle outside it. What about generation instead of multiple-choice — does the compact cache hold up for actual summarization? 32 00:08:18,403 --> 00:08:49,128 [Dr. Ada Shannon] It does — the more practical result, maybe. On HELMET's multi_lexsum, judged against full- and no-context anchors, Still recovers 74 to 95 percent of the full-context gain from 8k to 64k, still 59 percent at 128k — ahead of Attention Matching, Zweiger and colleagues out of MIT, 2026, and KV-Distill at every length. On LongBench summarization, GovReport and QMSum, it wins 300 of 500 head-to-head judge comparisons against KV-Distill. 33 00:08:49,128 --> 00:09:15,878 [Hal Turing] So — what actually worries me here. The entire training corpus is synthetic extractive multiple-choice, generated and verified by the same frozen model that later plays teacher. Doesn't that risk the compactor just learning to nail the right letter, rather than a genuinely general representation? And the introduction's whole motivating scenario is multi-day coding agents and tool use — neither shows up anywhere in training or evaluation. 34 00:09:15,878 --> 00:09:50,978 [Dr. Ada Shannon] It's a real risk, baked into the eval design — optimizing for 'pick the right letter' across four domains isn't the same as a robust general representation. O'Neill's got form probing this one level down too — he's also behind 'Can a Language Model Learn Facts Continually in Its Weights?', asking whether something durable is encoded or the model's just pattern-matching a familiar test shape. Still never answers that here, and there's zero agentic tool-use evaluation anywhere — no multi-turn tool calls, no long-horizon coding traces — despite that being the exact scenario the abstract opens with. 35 00:09:50,978 --> 00:10:16,478 [Hal Turing] Wait, hold on — that plugs right into something that bugged me on the summarization side too. Those HELMET and LongBench numbers aren't from the compactor we've been discussing all episode — the generation-aligned checkpoints get fifty extra fine-tuning steps on the actual target-task summary rows first. So is that really the same compact cache generalizing, or a fine-tuned descendant with homework for GovReport specifically? 36 00:10:16,478 --> 00:10:38,353 [Dr. Ada Shannon] That undercuts the single-general-compactor pitch more than the paper lets on. Needing a per-task fine-tune means you're deploying task-specific descendants sharing an architecture, not one compactor for anything a frozen model faces. Fifty steps is cheap, so the base representation clearly has useful structure — but 'the cache preserves the gain' and 'a fine-tuned descendant preserves it' are different claims, and only the first made the abstract. 37 00:10:38,353 --> 00:11:05,253 [Hal Turing] Speaking of what's in the abstract versus buried three appendices deep — Appendix B-one, the actual vLLM test. Single request, 64k context, one H200. Still's compact cache: six-point-seven-six seconds to first token. Full uncompressed cache: three-point-one-five. The compact path is slower end to end, and the abstract still calls this 'the favorable side of the speed-quality frontier.' 38 00:11:05,253 --> 00:11:25,653 [Dr. Ada Shannon] Because that frontier plot measures offline compaction time, not serving latency — different numbers. Burying the one real serving measurement in a single appendix table while the headline word is 'fast' is a framing choice I'd push back on. What you're buying today is retained memory, not a faster response, and there's no break-even point reported anywhere. 39 00:11:25,653 --> 00:11:48,878 [Hal Turing] I actually disagree with you there, Ada. They ran the honest microbenchmark and published the worse-case table themselves — nobody made them include it. That's more candor than most compaction papers bother with. Calling the memory reduction the frontier win isn't dishonest if the results section is upfront that this is compaction speed, not serving TTFT. 40 00:11:48,878 --> 00:12:09,203 [Dr. Ada Shannon] No, no — appendix candor doesn't fix abstract framing. Nobody reads Appendix B before deciding whether to prototype this. The abstract says 'favorable side of the speed-quality frontier,' full stop, no qualifier about which speed. A reader skimming it walks away thinking this makes inference faster. It doesn't, not in the one benchmark that actually measures it. 41 00:12:09,203 --> 00:12:20,029 [Hal Turing] Okay — fair. 'Honest three sections in' and 'clear upfront' aren't the same bar. I'll grant the abstract oversells speed even where the appendix redeems it. 42 00:12:20,029 --> 00:13:04,229 [Dr. Ada Shannon] And speaking of things left unquantified — Cartridges, Eyuboglu, Ehrlich, Arora, Guha and colleagues, 2025, the strongest quality reference here, never gets a number next to it. Cartridges optimizes the cache per context via prefix-tuning on self-generated dialogue; Still claims to be its amortized cousin, but without a head-to-head we can't judge the accuracy cost of that speed. Maurizio Diaz's 'Learned structure in cartridges,' 2025, found trained Cartridges use keys as stable retrieval routers while values carry the content — Still's own Appendix M-three finds the identical pattern, which is convincing cross-method evidence. Add the untouched limitations: iterative compaction collapsing past its trained horizon, a fixed ratio rather than a real budget, near-zero needle retrieval, and a fresh compactor needed per checkpoint. 43 00:13:04,229 --> 00:13:12,929 [Hal Turing] Which is the honest summary: real gains in a narrow lane, and real distance from the agentic future this paper opens with. 44 00:13:12,929 --> 00:13:55,254 [Dr. Ada Shannon] Exactly — and that gap is what tells you what to actually do with this today versus what to wait for. The honest answer: today's win is memory, not speed. A full cache runs around 144 kibibytes per token once you sum every layer and head, so one 128,000-token agent session can eat tens of gigabytes of GPU memory. Still's compact cache brings that down to tens of MiB regardless of the original length. That's a capacity story — more concurrent long-context sessions resident in the same footprint. It doesn't make any single request faster; we already saw the vLLM number — worse time-to-first-token than the full cache. So if you're capacity-bound serving thousands of long-lived agent sessions, this is useful today. If you're latency-bound on one request, it isn't yet. 45 00:13:55,254 --> 00:14:18,829 [Hal Turing] So the practitioners this matters to right now are the ones running fleets of long-horizon agents — multi-day coding assistants, tool-use loops that keep accumulating context — where the constraint is how many sessions you can keep alive at once, not how fast any one of them responds. Narrower than 'everyone doing long context,' but real. Where does the paper say this heads next? 46 00:14:18,829 --> 00:14:38,379 [Dr. Ada Shannon] A few honest next steps, and I'll give them credit for flagging what's missing. First, curriculum training toward much longer iterative horizons — checkpoints are trained at fixed 8k, 16k, 32k windows right now, and they explicitly gesture at pushing that curriculum out toward something like a million tokens of accumulated— 47 00:14:38,379 --> 00:14:56,579 [Hal Turing] Wait, hold on — a million tokens? That's not an incremental target, that's basically 'stop worrying about context length, period.' If they actually land that, this stops being a memory optimization and starts being the interface for how long-horizon agents keep state at all. 48 00:14:56,579 --> 00:15:16,854 [Dr. Ada Shannon] I actually disagree with the framing, Hal. That's a roadmap bullet, not a result — we already talked through why iterative compaction collapses below the no-context floor past its horizon. Scaling the curriculum further doesn't automatically fix that, it might just relocate the cliff. I want the collapse boundary characterized before I call a million tokens plausible. 49 00:15:16,854 --> 00:15:39,154 [Hal Turing] Sure, but every amortized method starts undertrained on the exact regime it eventually handles — that's not disqualifying, it's just where they currently are. I'm not telling you to trust the million-token number today. I'm saying the underlying direction, learn a reusable compactor instead of refitting per context, is the right one to keep funding. 50 00:15:39,154 --> 00:16:00,704 [Dr. Ada Shannon] Fine — sound direction, aspirational number, we can both live with that. Two other honest items: a constant-budget, slot-reuse recurrence variant instead of today's fixed-ratio growth, so memory stays flat across iterations; and production kernel support for the bias channel they disabled here. Critically, they call for a matched head-to-head against Cartridges — the comparison this whole conversation has been missing. 51 00:16:00,704 --> 00:16:32,054 [Hal Turing] Which loops back to the bigger pattern you flagged at the top of the episode — amortized variational inference, sparse autoencoders, and now this: swap a per-instance optimization for a trained encoder that does the job in one pass. Still is that same move, applied to the KV cache. So for listeners: this isn't a faster cache, it's a smaller one — real memory savings today, honest open problems, and a research direction, not yet a product, toward much longer horizons. Ada, thanks for digging through this one with me. 52 00:16:32,054 --> 00:16:35,629 [Dr. Ada Shannon] Always a pleasure, Hal. Thanks for listening, everyone.