1 00:00:01,000 --> 00:00:50,424 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs — Yirui Liu et al., eight authors, out of the Institute of Artificial Intelligence at China Telecom, TeleAI, together with Shanghai Jiao Tong University, Xi'an Jiaotong University, and University at Buffalo. It hit arXiv on July 31st, 2026. And Ada, here's the number that got me: this paper takes something the field assumed was mathematically settled — how you combine multiple cached pieces of context into one — and shows the textbook-correct answer actually wrecks one class of model, while a cheaper shortcut wins outright. 2 00:00:50,424 --> 00:01:28,150 [Dr. Ada Shannon] And it's not a marginal wreck. Under the 'correct' approach, one of their three test models loses more than half its usable quality in some configurations, while the supposedly sloppier shortcut holds up fine. That should make anyone building serving infrastructure nervous, because the systems instinct is always: if you can compute the exact thing, compute the exact thing. This paper is a case study in that instinct failing the moment you leave full attention behind — and it fails for a specific, traceable reason tied to how a newer family of recurrent layers propagates information forward. Worth unpacking properly before we get to the wreckage. 3 00:01:28,150 --> 00:01:48,375 [Hal Turing] Let's back up first, though, because this sits at the intersection of two things listeners may only half-know: position-independent caching, and hybrid Mamba-attention models. Ada, start with caching — I think most people picture it as just 'don't recompute what you already computed,' which isn't wrong, but it's not the whole story either. 4 00:01:48,375 --> 00:02:42,875 [Dr. Ada Shannon] Right, that's the baseline case, prefix caching. Every attention layer produces a Key and Value vector per token, stored in a KV cache so you're not redoing attention over the whole prefix each time. vLLM's PagedAttention, Kwon et al. out of UC Berkeley, SOSP 2023, and SGLang's RadixAttention, Zheng et al., NeurIPS 2024, both reuse a cached chunk only if it appears at the exact same position, behind the exact same preceding tokens, as when it was first cached. Reorder a conversation, or have a RAG system retrieve the same document but slot it third instead of first, and that's a miss — full recompute. Position-independent caching relaxes that: prefill a chunk in isolation, cache it under a content identifier instead of a position, then splice it into any request in any order, recomputing just a small selected slice of tokens afterward to patch up the cross-chunk attention that isolation lost. 5 00:02:42,875 --> 00:03:03,875 [Hal Turing] Okay, that tracks for a normal transformer, where every layer keeps this token-by-token ledger. But this paper is about hybrids swapping in Mamba-2 or Gated DeltaNet layers instead of full attention. My gut says caching a ledger and caching whatever those layers use internally are not remotely the same problem. 6 00:03:03,875 --> 00:03:47,000 [Dr. Ada Shannon] That's exactly the mismatch, and it's the whole reason this paper exists. Hybrid LLMs interleave a handful of full-attention layers with a majority of linear-recurrent layers — Mamba-2, Dao and Gu, ICML 2024, or Gated DeltaNet, Yang, Kautz, and Hatamizadeh, ICLR 2025. Those recurrent layers keep no token-indexed cache at all. They compress everything they've processed into one fixed-size hidden state, updated token by token — architecturally closer to an old RNN's hidden state than to a transformer's KV cache. Mamba-2 updates that state with a simple scalar decay, the same fade applied uniformly in every direction. Gated DeltaNet instead applies a dense, direction-dependent correction, so it can selectively preserve some parts of the state and overwrite others. 7 00:03:47,000 --> 00:04:01,275 [Hal Turing] Wait, wait, hold on — if it's just one blob of state per chunk, and PIC's whole trick was 'recompute a few tokens to patch the cache,' what is there left to patch? There's no per-token entries to selectively touch anymore. 8 00:04:01,275 --> 00:05:15,675 [Dr. Ada Shannon] Exactly the right question — that's challenge one in the paper. Full-attention PIC's core moves, concatenate the KV, selectively repair a few tokens, have no equivalent on the recurrent side. So the real question becomes: can you still get PIC's payoff — cutting time-to-first-token, the latency before a model produces its first output token, which is the entire point of caching — on hybrid models, and what does 'combining several cached states into one' even mean when there's no per-token concatenation to lean on. LinearKV's answer is what it calls decoupled initialization, applied right after chunk matching. Full-attention layers do exactly what they always did: concatenate the cached KV in context order. But each recurrent layer gets a separate rule, some function f that takes the K matched chunks' cached local states and folds them into a single initial state before recomputation even starts. And there turn out to be exactly two serious candidates for f. One is exact composition — literally compose all K cached states through their recurrence transitions to reconstruct the true full-prefix state algebraically. That's what the concurrent HYPIC paper, from Yifei Liu and co-authors in 2026, does. The other is LinearKV's own move — take just the single most recent matched chunk's cached state and use that alone, throwing the other K-minus-one away. 9 00:05:15,675 --> 00:05:36,574 [Hal Turing] Hold on — throwing away information sounds like it should just be strictly worse. Exact composition isn't an approximation of the algebra, it's the literal closed-form telescoping sum through the real transitions. If tossing K-minus-one states can somehow beat that, something's off with calling it 'exact' in the first place. 10 00:05:36,574 --> 00:06:21,774 [Dr. Ada Shannon] It's exact relative to the wrong ground truth, that's the catch. Each chunk's cached transition was built from a chunk prefilled in total isolation, blind to every earlier chunk. That's harmless at layer zero, where every chunk starts from the same empty history. But deeper layers read hidden inputs that, in a real joint prefill, already carry cross-chunk context — so composing exact algebra over these mismatched operators doesn't reproduce the truth. The paper frames it as a recursion: the error at chunk j equals the transition applied to the previous error, plus a freshly injected mismatch. Whether that error gets retained or damped depends entirely on the transition itself. Mamba-2's transition is a bare scalar decay — it can only retain and sum error forward. GDN's is a dense, gated rank-one correction — it can suppress or overwrite specific directions and actually hold the error down. 11 00:06:21,774 --> 00:06:34,049 [Hal Turing] Oh wait, wait — that predicts something you could just go measure directly, right? Track the error layer by layer and you'd expect it flat on GDN and blowing up somewhere deep on Mamba-2. 12 00:06:34,049 --> 00:07:06,324 [Dr. Ada Shannon] Exactly what they do — Figure 3. At layer zero the error sits near 0.01 on every model, basically the bf16 rounding floor, confirming the divergence at depth is real model behavior, not a bug. On the GDN models, OLMo and Qwen, the error stays under 1.0 at every single layer. On Granite, the Mamba-2 model, it compounds: one deep layer spikes to roughly 2x the state norm, stable across all five LongBench datasets, with several other layers sitting above 1.0 too. 13 00:07:06,324 --> 00:07:13,574 [Hal Turing] So does that error actually show up in the answers the model gives, or is it just a norm on paper? 14 00:07:13,574 --> 00:07:58,174 [Dr. Ada Shannon] It shows up. Setup is fixed 512-token chunks, offline-cached, and the same three plug-in selectors reused completely unchanged — CacheBlend, EPIC, ProphetKV. Only the initializer swaps. Under EPIC at a matched 20% recompute budget, exact composition on Granite recovers just 46.6% of full quality; single-block initialization recovers 86.8%. On the two GDN models it's a wash — the two methods tie, both climbing up to 92% of full depending on selector and benchmark. And single-block is cheaper too: it cuts time-to-first-token to about 0.46x of full prefill at 32K, while exact composition needs 5 to 17% more overhead on top of that, because online it has to fold a dense per-chunk transition matrix instead of just reading one cached state. 15 00:07:58,174 --> 00:08:04,249 [Hal Turing] Can't you just throw more recompute budget at exact composition until it catches up? 16 00:08:04,249 --> 00:08:43,375 [Dr. Ada Shannon] They tested that — swept the budget from 3% up to 40%, holding the selector and positions fixed, varying only the initializer. On Granite, exact composition stays flat at 41 to 52% of full quality at every single budget, while single-block reaches 76 to 89%. The gap never closes, even at 40% recompute — the initializer sets a quality ceiling that recompute can't buy back. And there's a control ruling out 'it's just that the last chunk is special': a randomly chosen single block does about as well as the last one, within roughly a point of Avg-F1 on Granite. So it's reading from a single source that matters, not which particular chunk. 17 00:08:43,375 --> 00:08:59,350 [Hal Turing] I want to flag something though — the whole 'Mamba-2 is fragile' headline rests on exactly one Mamba-2 model, Granite-4.0-H-Tiny, against two GDN models. That's not a controlled architecture comparison, that's an N of one. 18 00:08:59,350 --> 00:09:18,375 [Dr. Ada Shannon] Fair, and to their credit they say so themselves — the paper's own words are 'an empirical boundary over the evaluated models... not a proven law for every hybrid.' It's a reproducible effect on this model across five selectors, benchmarks, and 4x context, but it's still one data point on the Mamba-2 side. 19 00:09:18,375 --> 00:10:05,950 [Hal Turing] And here's what makes me even less comfortable trusting Granite as the reference Mamba-2 point: its own full-recompute score on RULER's common-word-extraction subtask is 0.003 — that's below its no-recomputation, do-nothing naive-reuse baseline of 0.192. The upper bound is worse than the floor. The paper calls that 'a model-level failure, not a reuse artifact,' and to be fair, they show OLMo and Qwen score 0.869 and 0.998 on that same subtask under the identical harness, so it's isolated to Granite. But if Granite is already producing degenerate output on one subtask under their own pipeline, how sure can we be it's a clean baseline everywhere else the comparison runs? 20 00:10:05,950 --> 00:10:44,900 [Dr. Ada Shannon] You can't be fully sure, and that's the honest answer. They do the right partial fix — excluding CWE bumps Granite's RULER-8K full-recompute average from 0.700 to 0.800, and they say the reported percent-of-full numbers only shift by up to 2.3 points either way, so it's not silently inflating the headline gap. But that's a correction for one known-bad subtask, not a certification that Granite behaves like a generic Mamba-2 model everywhere else. It's entirely plausible that whatever makes Granite degenerate on CWE — some MoE routing quirk, some checkpoint-specific brittleness — is correlated with, or even the same root cause as, why exact composition detonates on it specifically. 21 00:10:44,900 --> 00:11:04,600 [Hal Turing] Wait, hold on — that's actually a pretty different framing than 'Mamba-2 is fragile.' You're saying it could be 'this particular Granite checkpoint has some kind of instability, and exact composition just happens to be the operation that exposes it' — which is a much narrower and much less quotable claim. 22 00:11:04,600 --> 00:11:37,325 [Dr. Ada Shannon] Right, and I think that's the responsible way to report it until someone runs a second Mamba-2 hybrid through the same pipeline. Eq. 7, their error-recursion formula, is actually useful here because it's falsifiable — it predicts scalar-decay transitions retain and sum error while dense gated transitions can suppress it. That's a property of the recurrence family, testable on the next Mamba-2 hybrid that ships. Until then, 'Mamba-2 fragile' should be read as 'this one checkpoint's error accumulated in exactly the way the scalar-decay math predicts it might' — suggestive, not proven. 23 00:11:37,325 --> 00:11:46,575 [Hal Turing] So zooming out — this paper isn't the only group that noticed hybrid PIC needed a new answer. What's the story with HYPIC? 24 00:11:46,575 --> 00:12:29,950 [Dr. Ada Shannon] HYPIC — Yifei Liu, Juntong Wu, Yang Liu, and colleagues, 2026 — is concurrent work solving the exact same problem, and it picks exact composition, the mathematically principled answer, as its default. Two groups converge on the same question and land on opposite engineering bets: HYPIC trusts the algebra, LinearKV trusts the empirics. That's genuinely interesting as a field moment — it means the 'obviously correct' move wasn't obvious enough to be uncontested. Worth noting the comparison here isn't fully apples-to-apples either — LinearKV only borrows HYPIC's seam selector, not its boundary-token exclusion or causal-convolution warm-up, so they're stripping HYPIC down to isolate the initializer, which somewhat favors LinearKV by handicapping HYPIC's full pipeline. 25 00:12:29,950 --> 00:12:35,175 [Hal Turing] And Marconi's the other reference point here, right? How does that fit in? 26 00:12:35,175 --> 00:13:10,775 [Dr. Ada Shannon] Marconi — Rui Pan, Zhuang Wang, Zhen Jia, and Tri Dao's group, MLSys 2025 — is the prior state of the art for hybrid caching, but it only reuses recurrent state at exact prefix checkpoints. That's the old prefix-caching constraint, just ported to hybrids. LinearKV and HYPIC both go further, letting chunks be matched and reused in any order, which is the actual position-independent part. So the lineage is: Marconi solved prefix reuse for hybrids, and this generation of work — LinearKV and HYPIC simultaneously — is racing to solve position-independent reuse for hybrids. 27 00:13:10,775 --> 00:13:33,950 [Hal Turing] One thing that bugged me reading through this: chunk size is fixed at 512 tokens the entire time, inherited from ProphetKV's convention. But the whole argument for why exact composition compounds error is about how many operators get chained together. Doesn't the number of chunks matter a lot here? Twenty tiny chunks feels structurally different from two huge ones. 28 00:13:33,950 --> 00:14:08,550 [Dr. Ada Shannon] It should matter, and they never sweep it. More chunks means more composed operators for exact composition to chain, which by their own Eq. 7 logic should make the collapse worse, not better — so their headline gap might even be conservative. But it also matters for the 'random block works fine' result in the appendix — that holds at 512-token chunking with whatever K that produces per document. Push toward many small chunks and single-block might start throwing away too much; push toward few huge chunks and there's barely any K left to discard. It's a real gap in the design space, not just a missing ablation. 29 00:14:08,550 --> 00:14:15,725 [Hal Turing] Okay, let's land this. If I'm running a hybrid-model serving stack today, what do I actually take away? 30 00:14:15,725 --> 00:15:06,400 [Dr. Ada Shannon] Don't build the algebraically elegant version. Single cached-state initialization is cheaper on TTFT and, on at least one Mamba-2 model, dramatically more robust than composing everything exactly — simpler engineering beats mathematical purity here. Just hold the caveats: this is validated on static RAG documents, not the multi-turn conversations and agentic reordering the paper actually motivates itself with, and batched serving with real cache-transfer costs is explicitly future work, so the TTFT numbers are a best case. Funnily enough, Xuelong Li, the corresponding author here, was also behind that computation-bandwidth-memory trade-off paper — this fits the same worldview, that the right memory trade-off is often the cheap approximate one, not the exact one. 31 00:15:06,400 --> 00:15:20,625 [Hal Turing] So the takeaway: position-independent caching can be extended to hybrid models, one cached state really is enough, and the 'exact' answer only wins on paper. Ada, thanks for walking through this one with me. 32 00:15:20,625 --> 00:15:22,150 [Dr. Ada Shannon] Anytime, Hal. 33 00:15:22,150 --> 00:15:27,825 [Hal Turing] That's it for today's episode. Thanks for listening, and we'll catch you next time.