1 00:00:01,000 --> 00:00:52,524 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called "Training Transformers for KV Cache Compressibility," by Yoav Gelberg and Yam Eitan as equal-contribution lead authors, with Michael Bronstein, Yarin Gal, and Haggai Maron rounding out five co-authors total, out of the University of Oxford, Technion – Israel Institute of Technology, AITHYRA, and NVIDIA. It went up on arXiv on May 12th, 2026. And Ada, the thing that jumps out immediately is the claim in their abstract: they say almost any sequence-to-vector function a transformer can compute admits both a version that compresses down to almost nothing, and a version that's fundamentally incompressible, even though both compute the exact same function. 2 00:00:52,524 --> 00:01:31,150 [Dr. Ada Shannon] Right, and that's the sentence that should stop you mid-sip of coffee, because it means compressibility isn't something you discover about a model after the fact — it's something that got baked in, or not, during training, almost as an accident of which solution gradient descent happened to land on. Every serving team running long-context inference has been treating KV cache compressibility as a property of the input: 'this document is repetitive, so it compresses well,' or 'this conversation is dense, so it doesn't.' This paper is saying no, actually, two models can see the identical input and one of them hands you a cache you can crush to a tenth of its size, and the other one just can't, structurally, no matter how clever your compression algorithm is. 3 00:01:31,150 --> 00:01:46,075 [Hal Turing] Okay, so before we get into what they actually do about it, let's make sure everyone's on the same page about why the KV cache is even a problem worth a whole paper. Ada, walk me through it — what is this cache, mechanically? 4 00:01:46,075 --> 00:02:23,699 [Dr. Ada Shannon] Every time a transformer generates a token, it computes attention over everything that came before it. To avoid recomputing that from scratch at every single step, the model caches the key and value vectors it already produced, for every token, every layer, every attention head. That's the KV cache. The problem is it grows linearly with sequence length — double the context, double the memory, and double the work to scan through it at decode time. For a short chat, that's nothing. For a codebase-sized context, a long legal document, or an agent that's been running for hours, that cache can dwarf the size of the model weights themselves and become the actual bottleneck on how many users you can serve concurrently. 5 00:02:23,699 --> 00:02:28,924 [Hal Turing] So people have obviously tried to fix this. What's the landscape look like? 6 00:02:28,924 --> 00:03:25,074 [Dr. Ada Shannon] Two broad camps. One camp redesigns the architecture itself — linear attention mechanisms, state space models, sparse attention variants — so the cache never gets that big in the first place. Those save memory but historically trade away some raw performance at scale, which is part of why plain transformers still dominate frontier models. The other camp leaves the pretrained transformer completely alone and intervenes only at inference time, either heuristically — things like heavy-hitter eviction from the H2O paper, Zhang and colleagues, 2023, or attention-sink methods like StreamingLLM, Xiao and colleagues, also 2023 — or with actual optimization. That's where Cartridges comes in, from Eyuboglu and colleagues out of Stanford's Hazy Research group, and Attention Matching from Zweiger and colleagues, both of which run a small optimization procedure to squeeze an existing cache down while trying to preserve behavior. 7 00:03:25,074 --> 00:03:46,250 [Hal Turing] Wait, wait — hold on, that's actually the crux of it, isn't it? Every one of those inference-time methods you just named is working with whatever representations the model happened to learn. If the model's representations are secretly non-compressible, all of that optimization machinery is fighting a losing battle from the start. 8 00:03:46,250 --> 00:04:31,050 [Dr. Ada Shannon] Exactly, and that's the paper's whole pitch. They formalize this as "KV-compressibility" — a property of a specific transformer, not the task: does there exist some compression policy that can shrink its cache to a smaller budget while keeping output error below a threshold, across sequences up to some length. And they don't just assert this is possible in principle, they prove it. Their thought experiment is a histogram: imagine a function that just counts occurrences of tokens. You can build a transformer that solves this by writing the running count into a single slot and updating it in place — trivially compressible, one KV pair carries everything. Or you can build an equally correct transformer that spreads the count across every token's cache entry redundantly, so nothing is disposable. Same input, same output, wildly different compressibility, and both are legitimate transformer implementations of the identical function. 9 00:04:31,050 --> 00:04:50,900 [Hal Turing] I have to push back a little here, though — isn't that a bit of a toy example? Real language models aren't computing histograms, they're doing something enormously more complicated, so I'm not sure a hand-built worst case tells us much about what actually happens when you pretrain a model the normal way on real data. 10 00:04:50,900 --> 00:05:13,350 [Dr. Ada Shannon] I actually disagree with you there, Hal. The point of a worst-case construction isn't that models literally look like that — it's an existence proof that closes off the lazy assumption that compressibility just falls out for free from ordinary training. If it's possible to construct an incompressible-by-design transformer for something as trivial as counting, you can't wave your hands and say a model trained at scale will naturally land in the compressible regime. Something has to push it there. 11 00:05:13,350 --> 00:05:31,475 [Hal Turing] Okay, that's fair — and honestly that reframes the whole rest of the paper for me. If compressibility has to be actively pushed for rather than assumed, then the interesting question becomes: what's the mechanism that does the pushing? Which I think is exactly where you're about to take us. 12 00:05:31,475 --> 00:05:53,300 [Dr. Ada Shannon] It is. They call it KV-Compression-Aware Training, or KV-CAT — a continued pretraining procedure that deliberately masks out KV slots during training itself, forcing the model to solve the task using fewer active slots, so that by the time any post-hoc compressor like Cartridges or Attention Matching gets applied afterward, it's operating on a model that was already nudged toward giving it an easier job. 13 00:05:53,300 --> 00:06:34,949 [Dr. Ada Shannon] ...gets applied, there's already structured redundancy to exploit instead of the model fighting it the whole way. Mechanically, they take a pretrained checkpoint and insert learned routers between groups of consecutive layers. Each router looks at the token representations from the layer below and produces a scalar score between zero and one for every token — an importance estimate. That score gets thresholded into a binary mask: above the cutoff, the token's KV pair stays active in attention; below it, that slot goes dark for this forward pass. They're careful about initialization, too — every router starts with all scores at one, so at step zero the model behaves exactly like the dense original. Training then gradually teaches it which slots it can afford to drop. 14 00:06:34,949 --> 00:06:53,300 [Hal Turing] So the router's essentially a gate that gets slowly turned down over training. But that raises the obvious question, Ada — what's actually teaching it where to cut? You can't just tell a network 'be sparse' and hope accuracy holds together on its own. There has to be something shaping that behavior. 15 00:06:53,300 --> 00:07:30,050 [Dr. Ada Shannon] There are three loss terms pulling in different directions. The first is a self-distillation term: they run a masked forward pass and a dense forward pass on the same sequence, and train the masked distribution to match the dense one via KL divergence, with the dense side stop-gradiented as a fixed teacher. The second is a budget loss — borrowed from Hwang and colleagues' 2025 dynamic chunking work — pushing the router's average masking rate toward a fixed target retention rate, rho. Third is a plain next-token-prediction loss on the dense pass itself, keeping the uncompressed model honest while all this masking pressure is applied. All three get summed and backpropagated jointly. 16 00:07:30,050 --> 00:07:51,275 [Hal Turing] Okay, and scale-wise — what did they actually train this on, and just as important, what ships at the end? Because if the masked forward pass is this whole training-time fiction, I want to know whether a deployed model is actually running masked at inference, or whether that's purely a training-time trick that disappears once you deploy. 17 00:07:51,275 --> 00:08:33,500 [Dr. Ada Shannon] Straightforward on both counts. They continue-pretrain two QWEN2.5 checkpoints — the 0.5B and the 1.5B — on FineWeb-Edu, about 5.24 billion tokens, using four learned routers shared across four layer groups rather than one router per layer, keeping overhead small. As for what ships: at inference you use the unmasked, fully dense forward pass — masking only happens during training, as a kind of stress test reshaping the representations. On top of that dense model, you apply a standard post-hoc compressor exactly as you would on any other checkpoint — Attention Matching, or gradient-based optimization in the style of Cartridges. KV-CAT doesn't touch the compression algorithm. It just hands that algorithm a friendlier model to work with. 18 00:08:33,500 --> 00:08:46,125 [Hal Turing] Oh — wait, hold on, sorry to jump in, but I need to know if this actually worked before we go any further. Did compression get better, or did they just make the uncompressed model worse to get there? 19 00:08:46,125 --> 00:09:27,625 [Dr. Ada Shannon] Both questions answered by the same table. First, zero compression applied, just running straight — six multiple-choice benchmarks, HellaSwag, WinoGrande, PIQA and others. KV-CAT lands within about half a point to seven-tenths of a point of the base model either direction — essentially a wash. Then, under a fixed optimization budget, using seven-hundred-sixty-eight-token prefixes reconstructing two-hundred-fifty-six-token suffixes across keep ratios from five to fifty percent, KV-CAT retains suffix perplexity up to three-point-two-one times better than the base model at matched settings. Letting gradient-based optimization run over many steps instead, KV-CAT hits the same perplexity target using up to five times fewer steps. 20 00:09:27,625 --> 00:09:48,625 [Hal Turing] I'll push back a little there, Ada — 'up to three-point-two-one times' is a headline number, and headline numbers in papers are almost always the best case in the table, not the typical case. Before I get excited, I want to know if that's an outlier at one keep ratio and one model size, or if it actually holds up across the board. 21 00:09:48,625 --> 00:10:08,925 [Dr. Ada Shannon] Fair challenge, but it holds up — not a cherry-picked cell. Table 3 reports it across four keep ratios, five, ten, twenty, forty percent, on both model sizes, and KV-CAT beats the base model on every one, by KL divergence and top-1 agreement too, not just perplexity. Three-point-two-one is the ceiling, but the floor in that table is still a meaningful improvement, not a rounding error. 22 00:10:08,925 --> 00:10:31,751 [Hal Turing] Okay, that's fair — consistency across the table is a much stronger claim than one flashy number, I'll grant you that. So beyond perplexity retention, does any of this translate into something a listener would actually notice, like retrieval or real question answering? Because perplexity is a little abstract for anyone who isn't staring at loss curves. 23 00:10:31,751 --> 00:11:12,301 [Dr. Ada Shannon] It does, on both fronts. A needle-in-a-haystack test — compress the haystack, then check exact-match retrieval of a planted passkey — and KV-CAT lifts mean retrieval accuracy by six-point-four points on the 0.5B model and five-point-two on the 1.5B, with the biggest jumps, eleven to nineteen points, in that thirty-to-fifty-percent keep-ratio range. On LongBench v2, seven long-context QA subdomains truncated to thirty-two-thousand-seven-hundred-sixty-eight tokens via middle truncation, compressed with the gradient-based method, KV-CAT improves average accuracy across every retention level, up to thirty-nine percent. Real tasks, not just a perplexity chart. 24 00:11:12,301 --> 00:11:32,651 [Dr. Ada Shannon] right, at the higher keep ratios specifically. So stack that next to the perplexity and LongBench numbers we already covered, and it's the same shape of improvement across three different evaluation setups — genuinely the strongest methodological point in the paper. But consistency under their conditions isn't the same as generalizing past them, and that's where I get more skeptical. 25 00:11:32,651 --> 00:12:09,176 [Hal Turing] Which conditions, specifically? Every number in Tables 1 through 4 comes from Qwen2.5-0.5B and 1.5B. The paper opens by framing the KV cache bottleneck as a serving problem — one that gets brutal at seven billion, seventy billion parameters, well past anything tested here. Their own Limitations section admits it 'remains to be seen how well the method generalizes to larger-scale models.' So how much can sub-2B continued-pretraining experiments really tell us about whether this survives where the bottleneck actually bites? 26 00:12:09,176 --> 00:12:53,526 [Dr. Ada Shannon] Not a lot on its own, and it compounds with sequence length: training and most evaluation happens at exactly 1024 tokens, a 768-token prefix plus 256-token suffix. Even LongBench v2, their long-context showcase, truncates contexts to 32,768 tokens via middle truncation before compression starts — and 'long-context' in production means thirty-two K up to a million-plus. A bias learned at 1024 tokens could just be a short-sequence artifact. Third wrinkle: the budget loss trains the router toward one fixed retention rate, borrowed from Hwang, Wang, and Gu's dynamic chunking work at Carnegie Mellon, but evaluation sweeps down to five percent keep — ten times below whatever it was tuned for. At that point I'm not sure we're measuring the router's bias or the optimizer just compensating. 27 00:12:53,526 --> 00:13:32,201 [Hal Turing] Oh — wait, hold on, that ties into something else: both compressors they test, Attention Matching and the gradient-based Cartridges-style method, are expensive per-example optimization procedures. The cheap, training-free stuff that's actually deployed — H2O, SnapKV, StreamingLLM, PyramidKV — gets cited but never tested. And on LongBench, KV-CAT loses to the base model on many-shot learning at ten percent keep, nineteen versus twenty-eight-six, off maybe thirty examples per subdomain. Is 'up to thirty-nine percent' representative, or two favorable cells doing the work? 28 00:13:32,201 --> 00:14:35,326 [Dr. Ada Shannon] Probably some of both — no variance reported, small samples, so the headline max is exactly the number I'd trust least. The heuristics gap is real: KV-CAT is explicitly orthogonal to the compressor, reshaping the model rather than proposing an algorithm — same relationship to Cartridges, Sabri Eyuboglu and coauthors out of Stanford, 2025, and Attention Matching, Adam Zweiger, Xinghong Fu, Han Guo, and Yoon Kim at MIT, 2026. Lexico — Junhyuck Kim, Jongho Park, Jaewoong Cho, and Dimitris Papailiopoulos — does sparse dictionary coding, cited but never tested as a target. DuoAttention, Xiao and Song Han's group at MIT, takes a different angle, picking which attention heads need full retrieval at the head level instead of masking slots during training. None of this touches Theorem 3.1 either — an existence proof about hand-constructed weights designed to be maximally incompressible; whether ordinary SGD pretraining drifts toward that construction is asserted, not shown. Side note — Haggai Maron here also co-authored Nemotron 3 Ultra, the hybrid Mamba-transformer reasoning model out of NVIDIA. 29 00:14:35,326 --> 00:14:58,526 [Hal Turing] Okay, but I want to push back on how hard we're leaning on this. It's continued pretraining, a proof of concept — nobody's first paper validates at seventy billion parameters, that's compute most labs don't have. The theory and the recipe are architecture-agnostic. Isn't it a little unfair to dock them for not running an experiment that costs more than their whole compute budget? 30 00:14:58,526 --> 00:15:28,151 [Dr. Ada Shannon] No, I actually disagree with you there, Hal. It's not about begrudging them a cluster. The title is 'Training Transformers for KV Cache Compressibility' — framed as a general recipe, not a pilot. And there's a control they could've run at zero extra budget: continue-pretrain the base model on the same five-point-two-four billion FineWeb-Edu tokens, no masking, no budget loss, and see how much of the gain is just more training versus the objective itself. That's not more compute — it's the same run, minus two loss terms. Skipping it is a choice about what to report. 31 00:15:28,151 --> 00:15:51,401 [Hal Turing] ...Yeah, okay, that lands. So practically — where does this leave a team building on it? Worth prototyping at your own scale before trusting the headline numbers, worth a look if you're already running Attention Matching or Cartridges-style compression, but not something to bet a serving stack on before someone reruns it above two billion parameters and past 1024 tokens. 32 00:15:51,401 --> 00:16:13,826 [Dr. Ada Shannon] Add to that: nobody's tested whether this composes with quantization, the compression axis actually shipping in production today, or how it behaves under cross-request prefix caching, where a token the router calls disposable might get revisited turns later in an agent session. Scaling past 7B, training at real long-context lengths, and checking whether one checkpoint generalizes across many keep ratios instead of the one it was tuned for — that's the next paper. 33 00:16:13,826 --> 00:16:35,276 [Hal Turing] So the honest read: solid theory, a recipe that's consistently better than the base model at the scale and lengths they actually tested, and a self-acknowledged gap between that and the production long-context story the paper opens with. Worth watching, not worth deploying yet. That's it for this one — thanks for listening, and see you next time.