1 00:00:01,000 --> 00:00:45,303 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into "Comparative Characterization of KV Cache Management Strategies for LLM Inference," by Oteo Mamo et al. — four authors total, with Olga Kogiou, Hyunjin Yi, and Weikuan Yu — out of Florida State University, posted to arXiv on April 6th, 2026. And Ada, here's the number that got me: baseline H2O and InfiniGen both hit out-of-memory errors at around 10,000 tokens. That's less than ten percent of the 128K context window these models are supposed to support. 2 00:00:45,303 --> 00:01:23,291 [Dr. Ada Shannon] Right, and that's the number that reframes the whole paper for me. Everyone assumes the KV cache fight is about how clever your eviction policy is. This paper says: forget eviction policy, half these frameworks can't even survive prefill on a long document. No one had actually lined up vLLM, H2O, and InfiniGen side by side under identical hardware and identical models before. Prior comparisons stayed inside one family — sparsification papers benchmarked against other sparsification papers. This is the first time someone dragged all three paradigms onto the same four-H100 node and made them fight. 3 00:01:23,291 --> 00:01:33,554 [Hal Turing] So walk me through what 'three paradigms' even means here, because I think some listeners hear 'KV cache management' and assume it's all the same idea with different knobs. 4 00:01:33,554 --> 00:02:16,325 [Dr. Ada Shannon] It's genuinely three different philosophies. First, memory management — that's vLLM, from Kwon et al. at UC Berkeley, SOSP 2023. It keeps the entire KV cache, doesn't throw anything away, it just organizes storage better with PagedAttention. Second, static sparsification — that's H2O, Heavy-Hitter Oracle, Zhang et al., NeurIPS 2023. It permanently deletes low-importance tokens during prefill and never looks back. Third, dynamic selection — InfiniGen, Lee et al., OSDI 2024. It keeps everything, just moves it to CPU memory and speculatively fetches back only what it thinks it'll need at each step. Same problem, three completely different bets on what to sacrifice. 5 00:02:16,325 --> 00:02:27,750 [Hal Turing] Okay, wait, back up for a second, because I want to make sure everyone's with us on why the cache exists at all. Why can't the model just recompute attention over the whole conversation every time it generates a new token? 6 00:02:27,750 --> 00:03:14,607 [Dr. Ada Shannon] Because that's quadratic cost you'd be paying on every single step. Transformers, going back to Vaswani et al. at Google, 'Attention Is All You Need,' NeurIPS 2017, compute attention by comparing every token against every other token. During prefill, you do that once for the whole prompt and stash the resulting key and value tensors — that's the KV cache. During decode, instead of recomputing attention over the entire history, you just reuse those stored keys and values and append the new token's contribution. That turns per-token generation cost from quadratic back to linear. The catch is the cache itself grows multilinearly — with sequence length, with layer depth, with the number of attention heads. 7 00:03:14,607 --> 00:03:26,635 [Hal Turing] Oh wait wait wait — that's actually the part I wanted to push on, because doesn't Grouped Query Attention already fix a chunk of that? Llama 3.1 uses GQA, right, sharing key-value heads across query heads? 8 00:03:26,635 --> 00:03:57,750 [Dr. Ada Shannon] It does — Ainslie et al., Google Research, 2023 introduced GQA, and Llama-3.1-8B, from Dubey et al. at Meta, 2024, uses it to go from 32 query heads down to just 8 KV heads, a four-times reduction in cache size. That's real. But it's a fixed multiplier, and context windows didn't stay fixed. We went from 2 to 4K tokens a couple years ago to 128K now. A four-times cut against a workload that grew thirty-plus times over just gets swallowed. 9 00:03:57,750 --> 00:04:16,419 [Hal Turing] Right, but that also feels like an argument that this whole problem mostly resolves itself as hardware and architecture improve — bigger heads-savings, bigger GPUs, and eventually the cache just isn't the bottleneck anymore. Do we really need three separate management paradigms if the trend line is already bending the right way? 10 00:04:16,419 --> 00:04:36,713 [Dr. Ada Shannon] I actually disagree with you there, Hal. The trend line isn't bending fast enough, and the paper's own numbers prove it — GQA gives you a one-time four-times cut, but that's a constant, and sequence length growth is not bounded the same way. You can't architecture your way out of a workload that keeps expanding into six-figure token counts for every code repo or legal contract someone wants summarized. 11 00:04:36,713 --> 00:04:48,045 [Hal Turing] No no no, I hear that, but if the trend has held for two years running, isn't betting against it a little premature? Bigger context, sure, but hardware and head-sharing tricks have kept pace before. 12 00:04:48,045 --> 00:05:26,543 [Dr. Ada Shannon] They've kept pace on paper parameters, not on prefill memory, and that's exactly the gap this study exposes — GQA touches the cache's steady-state size, not the transient attention-score matrix you have to materialize just to decide what to evict in the first place. That's a structurally different cost, and it's why H2O and InfiniGen face-plant at 10K tokens even with GQA already baked into the model. So okay, we can agree to disagree on the long-run trend — but empirically, today, on this hardware, the three-paradigm question is very much alive, which is exactly why this study exists. 13 00:05:26,543 --> 00:06:18,092 [Dr. Ada Shannon] Constraint than head-sharing solves, because GQA shrinks steady-state cache size, but prefill still has to materialize the full attention-score matrix just to decide what's worth keeping in the first place. So here's what they actually built to test that: four H100s, Llama-3.1-8B and 70B, GPT-OSS-20B, LMSYS-Chat-1M for realistic prompts, a Wikipedia stream for stress-testing extreme lengths, six lm_eval reasoning benchmarks for accuracy. One calibration choice worth flagging up front: InfiniGen runs at a KV budget of 0.3, H2O at a Heavy-Hitter ratio of 0.3, implemented from scratch on top of FlexGen following the original methodology. Both were picked because at those settings each lands within 10% of baseline vLLM accuracy. 14 00:06:18,092 --> 00:06:43,030 [Hal Turing] Hold on — 'implemented from scratch on FlexGen' is the part I keep snagging on. H2O's own paper ships its own codebase. If FSU rebuilt the Heavy-Hitter selection logic themselves on a different inference engine instead of running the authors' actual implementation, how do we know a bad accuracy number later is the algorithm failing, versus their reimplementation just falling short of the real thing? 15 00:06:43,030 --> 00:07:33,417 [Dr. Ada Shannon] Fair, and the paper doesn't fully resolve it — FlexGen is Sheng et al.'s single-GPU throughput engine out of Stanford and Berkeley, ICML 2023, a reasonable substrate but not H2O's native harness, with no run against the authors' code to calibrate against. Keep that in your pocket. Now, time-to-first-token: vLLM scales cleanly to the full 128K window. H2O and InfiniGen, baseline configs, both OOM around 10,000 tokens — under ten percent of what the model supports. And it gets worse on paper, not better: that OOM boundary shows up earlier on multi-GPU setups, because parameters get sharded across devices but attention-score computation stays localized to one GPU. 16 00:07:33,417 --> 00:07:48,417 [Hal Turing] Wait, wait — hold on, that's completely backwards from what I'd have guessed. You add GPUs, you'd think you get more headroom, and instead the ceiling drops? Okay, so what actually fixes that — I assume that's where FlashAttention and Chunked Prefill come in? 17 00:07:48,417 --> 00:08:42,427 [Dr. Ada Shannon] Exactly, and they land very differently. FlashAttention-2 fuses attention into SRAM tiles so the full score matrix never materializes, dropping memory complexity to linear. For InfiniGen that's transformative — at 15,000 tokens, where baseline was already failing, FA cuts prefill memory 85 percent and latency 77 percent, and scales it all the way to the full 128K context. For H2O, FA does basically nothing, because Heavy-Hitter selection needs the actual materialized softmax of QK-transpose to score token importance — computing attention faster doesn't help if you still have to store the whole matrix to rank it. And I'd flag: that InfiniGen win is shown on Llama-3.1-8B, one GPU, 15K tokens. Nobody's shown it holds at 70B across four GPUs. 18 00:08:42,427 --> 00:08:55,941 [Hal Turing] Okay but if FA basically fixes InfiniGen's long-context problem, that feels like it closes most of the gap this paper opened with. Same kernel, same math, just a bigger matrix — I'd bet on it generalizing. 19 00:08:55,941 --> 00:09:22,737 [Dr. Ada Shannon] I actually disagree with you there, Hal. We just established OOM hits earlier specifically once you shard across devices, because attention stays localized to one GPU no matter how many others are in the box. Nobody's running 70B InfiniGen on a single H100 in production. Until someone shows this at 70B on four GPUs, I'm not calling the long-context problem solved — I'm calling it solved in the one configuration they happened to test. 20 00:09:22,737 --> 00:09:32,210 [Hal Turing] That's... actually hard to argue with, given the OOM finding. Fine, I'll take the caution flag. Still want the 70B number before I fully believe it. 21 00:09:32,210 --> 00:10:55,338 [Dr. Ada Shannon] Good, hold onto that skepticism. Speaking of asterisks — H2O's own workaround, Chunked Prefill, introduces a real algorithmic bug: tokens in different chunks never attend to each other during prefill, so Heavy-Hitter selection runs on incomplete information and evicts tokens purely for landing in the wrong chunk. On resource footprint, batch 16 to 96 on the 8B model: vLLM sits flat around 72 gigabytes GPU regardless of batch. H2O stays under 40 gigabytes even at batch 96, roughly half of vLLM. InfiniGen just relocates the burden — similar GPU footprint to H2O, but CPU memory blows past 100 gigabytes at batch 96, two-and-a-half times either alternative. Throughput matches: vLLM and H2O scale linearly and land in the same range, H2O actually about three times more throughput per gigabyte once normalized. InfiniGen is an order of magnitude slower throughout, because every decode step needs a serialized CPU-to-GPU transfer. Push output to 8,192 tokens and it compounds — InfiniGen needs over sixteen minutes where vLLM takes about one, H2O sits in the middle around three. 22 00:10:55,338 --> 00:11:06,808 [Hal Turing] Right, throughput and memory feel settled. What about accuracy — presumably where H2O's aggressive eviction actually shows its cost, especially with that reimplementation question still hanging over it. 23 00:11:06,808 --> 00:12:01,004 [Dr. Ada Shannon] At budget 0.5, both recover close to baseline, no story there. Drop to 0.1 and they diverge hard — InfiniGen degrades gracefully, about six points on average, while H2O falls off a cliff: HellaSwag down 23 points, BoolQ down 16, exactly the numbers I flagged earlier. Then the retention task — inject a fact early in a multi-turn conversation, ask about it later. vLLM holds 94 to 96 percent, InfiniGen tracks close at 92 to 94. H2O collapses to 70 to 78 percent, and it gets worse as context grows: 17 points below baseline at one-to-two-thousand tokens, 27 points below at four-to-eight-thousand. Permanent eviction doesn't just lose accuracy, it specifically loses the early stuff — exactly what you'd expect from a policy that deletes tokens and never looks back. 24 00:12:01,004 --> 00:12:44,147 [Dr. Ada Shannon] ...inject a verifiable fact — 'my dog's name is Barkley' — right at the start of a conversation and ask about it six or eight turns later. vLLM holds 94 to 96 percent recall as the baseline. InfiniGen tracks within a few points, 92 to 94. H2O collapses to 70 to 78 percent, and it gets worse the longer the context runs — 17 points below baseline at 1 to 2K tokens, 27 points down by 4 to 8K. That's basically Xiao et al.'s attention-sink insight from StreamingLLM, MIT, 2023, showing up by omission — H2O doesn't protect the sink tokens the way StreamingLLM does, so early facts just vanish. 25 00:12:44,147 --> 00:13:18,419 [Hal Turing] Here's what bugs me about matching them on aggregate accuracy in the first place, Ada. They tuned H2O and InfiniGen so both land within 10 percent of vLLM on the six reasoning benchmarks, called that 'equivalent,' and ran the whole system comparison on top of it. But the retention task uses that exact same budget of 0.3, and it shows H2O in the 70s against InfiniGen's 92 to 94. So the two were never really accuracy-matched — they just happened to average out the same across six benchmarks that don't touch long-range retention at all. 26 00:13:18,419 --> 00:13:55,710 [Dr. Ada Shannon] Right, and that 0.3 wasn't chosen blind — HH ratio and KV budget both got picked specifically because that's the point where each lands within 10 percent of baseline. That's one operating point on a curve, selected after looking at the accuracy results, then used to justify every throughput and memory claim downstream. We already flagged the FlexGen reimplementation as one confound; this is a second one stacked on top. Pick a friendlier budget, say 0.5, and H2O's memory advantage probably shrinks right along with its accuracy gap. Nobody characterized the full accuracy-memory surface, just one convenient slice of it. 27 00:13:55,710 --> 00:14:31,005 [Hal Turing] Oh, hold on — that same 'one convenient slice' problem applies to the FlashAttention result too, doesn't it? That 85 percent memory cut and 77 percent latency cut for InfiniGen is shown on Llama-3.1-8B, 15K tokens, single GPU — exactly the configuration where we don't yet know if the multi-GPU ceiling kicks in. Does 'near-linear to 128K' actually hold at 70B, or is that one small, favorable case standing in for the whole claim? 28 00:14:31,005 --> 00:14:58,172 [Dr. Ada Shannon] I'll push back on how hard you're leaning into that, Hal. The TTFT figure does include 70B and GPT-OSS-20B, and the OOM pattern holds at those scales too, so the multi-GPU ceiling itself is verified, not just assumed. I think it's fair to say the qualitative story generalizes even if the fine-grained FA percentages weren't re-measured beyond 8B. The paper isn't claiming precision at scale here, just direction. 29 00:14:58,172 --> 00:15:32,909 [Hal Turing] No, I actually disagree with you there, Ada. TTFT only tells you where prefill falls over — it says nothing about whether H2O stays three times more memory-efficient per gigabyte at 70B, or whether InfiniGen's transfer serialization gets worse once you're moving bigger KV tensors over the interconnect. Those are the headline numbers practitioners will actually act on — batch scaling, throughput, the 17x decode slowdown — and every one comes from a single 8B model on one GPU. Calling that a comparative characterization 'across model sizes' is a real overreach. 30 00:15:32,909 --> 00:16:05,417 [Dr. Ada Shannon] Fair — that's a legitimate gap, and I'll concede the practitioner guidance reads more universal than the evidence actually supports. Same caution applies to the NVLink-C2C line at the end, honestly. They say 7x PCIe bandwidth 'could substantially narrow' InfiniGen's gap, but that's speculative, and the bottleneck they describe is serialization — one transfer blocking the next decode step — not raw bandwidth. Faster links shrink each transfer's latency, sure, but if the dependency chain stays sequential, you're still paying transfer count times per-transfer cost. 31 00:16:05,417 --> 00:16:40,015 [Hal Turing] Good catch, and it's not the only thing left unexamined — nobody asks what happens to evicted tokens afterward. H2O just discards them, but there's a whole 2024-2026 line on reusing exactly that kind of KV — recompression, cross-request sharing. Weikuan Yu's on this paper and also on DUAL-BLADE, the NVMe-direct KV offload work for edge inference, so there's clearly appetite in this group for treating KV as reusable state — just not applied here. 32 00:16:40,015 --> 00:17:22,368 [Dr. Ada Shannon] Which is really the honest summary of the whole paper: no framework wins outright. vLLM when memory isn't the constraint and you want flat latency. H2O when you're memory-starved but need throughput and can tolerate the occasional dropped fact. InfiniGen when retention matters more than speed and you can eat the cost. The real contribution isn't a winner, it's turning 'which cache strategy' into an explicit three-way tradeoff instead of a default pick. A rigorous follow-up needs the full accuracy-memory curve per framework, H2O's actual codebase, and this whole battery repeated at 70B before any of it graduates from an 8B result to general guidance. 33 00:17:22,368 --> 00:17:42,012 [Hal Turing] Good summary. So: cache management comes down to memory versus throughput versus accuracy — pick your poison based on what you can least afford to lose, and don't trust a single-model number as gospel until someone reruns it at scale. Thanks for listening, everyone — that's all for this one.