1 00:00:01,000 --> 00:00:43,850 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper is Hypic: Accelerating Hybrid-Attention LLM Serving with Position-Independent Caching. First author Yifei Liu, et al., seven co-authors in all, out of Xiaohongshu Inc., Peking University, and Shanghai Jiao Tong University. Posted to arXiv July 12th, 2026. This is our KV-cache territory again, Ada — same ground as Breaking the Prefix Barrier with Shared KV Cache and TokenDance for Multi-Agent KV Cache Sharing, but a genuinely new wrinkle. 2 00:00:43,850 --> 00:01:16,150 [Dr. Ada Shannon] Here's the headline: this is, they claim, the first system that lets you cache independent prompt chunks on models that don't keep a normal per-token KV cache at all. A lot of production serving stacks now compress most layers into one fixed-size blob per request instead of a token-by-token cache, and every existing position-independent caching trick assumes that per-token cache exists. What pulled me in was the rigor — four hybrid-attention models, five workloads, and they isolate exactly where the error comes from instead of hiding behind one aggregate number. 3 00:01:16,150 --> 00:01:36,275 [Hal Turing] Okay, genuinely, walk me through this before we go further. Why does prefill, the one-time pass over the prompt before generation starts, dominate the cost? I always figured decoding, one token at a time, was the expensive part. Especially since the paper keeps framing this around RAG and agent prompts. 4 00:01:36,275 --> 00:02:18,574 [Dr. Ada Shannon] In RAG and agent pipelines the prompt itself is huge before generation starts — retrieved documents, tool descriptions, memory snippets, tens of thousands of tokens stitched together. Decoding one token is cheap; attending over fifty thousand tokens once, at the start, is not. That one-time pass is prefill, and TTFT, time-to-first-token, is basically how long it takes. Serving systems already cache KV for an exact literal shared prefix — that's position-dependent caching, what vLLM's Automatic Prefix Caching and SGLang's RadixAttention ship today. Position-independent caching goes further: cache each segment once, splice it behind any prefix. The idea traces back to the Prompt Cache paper, In Gim and Guojun Chen out of Yale, 2024. 5 00:02:18,574 --> 00:02:41,549 [Hal Turing] Right, and that's a great trick for RAG — same documents, different order every time. So where's the trap? You said some models don't keep per-token KV at all, that's this hybrid-attention thing? Production models like Qwen3.5, Ring-2.5, and MiniMax-M1 now swap most layers, three-quarters or more, over to some other kind of attention that works completely differ— 6 00:02:41,549 --> 00:03:14,125 [Dr. Ada Shannon] Oh, sorry, jump in, this is the part people flip. Linear attention doesn't keep every token's key and value around forever, growing with each token like a transformer normally does; it compresses the whole history into one fixed-size matrix, updated by a running recurrence, one token at a time. Same conceptual move as an RNN's hidden state, one blob summarizing everything you've seen, except these are built to train in parallel like a transformer, not sequentially. Cheap, constant memory. But there's no per-token handle into that state anymore — you can't point at 'token 47's key and value' the way you can with real KV cache. 7 00:03:14,125 --> 00:03:40,824 [Hal Turing] So that's the collision. Splice and correction, the two PIC primitives, both need a list of per-token vectors you can cut and paste, plus selectively recompute the few tokens that look wrong. A segment — some independent reusable chunk, a retrieved document, an agent's memory snippet — that only exists as one compressed blob gives you no list to cut. It's not that the trick gets harder; the primitive it's built on doesn't exist in these layers. 8 00:03:40,824 --> 00:03:58,099 [Dr. Ada Shannon] Right, and to their credit, they don't force the old splice trick onto a state that can't be spliced. They go find a different algebraic object the compressed state actually supports. That's the harder, correct move, not a patch. What that object actually genuinely is, we'll dig into next time. 9 00:03:58,099 --> 00:04:19,149 [Hal Turing] Okay, but push back for a second — nobody's built hybrid-attention PIC before this, not vLLM, not SGLang. Doesn't that suggest this is still mostly a theoretical problem? Are enough people actually running hybrid-attention models at real scale for this to matter, or is this solving a gap that's still mostly empty? 10 00:04:19,149 --> 00:04:49,074 [Dr. Ada Shannon] I actually disagree with you there, Hal. That's not a demand problem, it's a timing problem. Hybrid-attention stacks only reached real production maturity very recently — MiniMax-M1, Qwen3.5, Kimi-Linear, the Ring and Ling line — these are live shipping systems, not research toys. Non-contiguous PIC took years to become standard even for plain transformers, because it needs careful error bounding most teams never bothered with. The research simply hasn't caught up to the architecture shift yet. That's not 'nobody needs it' — that's 'this paper got here first.' 11 00:04:49,074 --> 00:05:06,899 [Hal Turing] No, that's fair, I'll take it. First-mover on a real architecture shift is a different claim than solving a made-up problem. I still want to see who actually deploys this in production before I call it a win, but the timing argument holds up better than I expected it to. 12 00:05:06,899 --> 00:05:23,349 [Dr. Ada Shannon] Good, hold that thought, because next up is the actual mechanism: the operator that lets two independently-cached states combine as if they'd been computed back to back, and it's stranger, and more elegant, than just recomputing everything from scratch. Trust me on this one. 13 00:05:23,349 --> 00:05:58,775 [Dr. Ada Shannon] Right, so here's the elegant part. Every segment, once it's cached, carries a transition operator, not just its raw end-state — the function describing how that segment would transform any state arriving right before it. The tempting shortcut is to just add two segments' end-states and assume you get what you'd have from running them back to back. Their Table 2 shows that fails even in the simplest family. Take RetNet, the Retentive Network line from Sun and colleagues at Microsoft Research, 2023. Add two RetNet end-states directly and the decay terms don't compose — you get a structurally different object, not a slightly noisy one. 14 00:05:58,775 --> 00:06:17,000 [Hal Turing] Okay, that's a bigger deal than I gave it credit for. So this isn't "our approximation has error bars," it's "the operation you'd naturally reach for doesn't even live in the same algebraic universe as the right answer." What's the fix — if you can't add states, what do you do instead? 15 00:06:17,000 --> 00:06:58,050 [Dr. Ada Shannon] You compose the operators instead of adding the states — chain them the way you'd multiply matrices rather than add them. That's the composition law: take segment one's end-state, left-multiply by segment two's transition operator, add segment two's own end-state, and the result is provably identical to running both back-to-back from scratch. It generalizes across Table 1's three linear-attention families — scalar decay like RetNet, a data-dependent diagonal like GLA, and dense-matrix transitions with a delta-rule erasure term. All three are affine operators, next-state equals a transition matrix times current-state plus a write term, they just differ in what that transition matrix looks like — scalar, diagonal, or full dense. 16 00:06:58,050 --> 00:07:10,450 [Hal Turing] Oh — wait, hold on, sorry to jump in — dense meaning a full matrix has to be way more expensive to carry around per segment than one scalar decay number, right? That can't be free. 17 00:07:10,450 --> 00:07:40,575 [Dr. Ada Shannon] It's not, no — Table 3 lays out exactly that gradient. Scalar-family transitions store as a couple of bytes per segment, diagonal costs a few hundred bytes, dense costs a full matrix, tens of kilobytes at typical head dimensions. But even the dense case is still nowhere near storing per-token KV for that segment — you're trading a cost fixed by matrix dimensions against a cost that scales with sequence length. The ordering holds throughout: composing operators, even the expensive dense kind, stays far cheaper than the thing PIC is replacing. 18 00:07:40,575 --> 00:07:55,975 [Hal Turing] Okay so state composition's solved. But splice-and-correction on the full-attention side always needed a correction step to fix token-level errors at the seam. If there's no per-token vector anymore, how do you correct anything? 19 00:07:55,975 --> 00:08:29,325 [Dr. Ada Shannon] That's the second wall. Full-attention correction works because it exposes a per-token hidden state you can selectively recompute. Linear layers don't give you that — you only ever see the compressed end-state, there's no token-indexed handle to patch. What they found instead is that the error from stitching two segments together isn't spread evenly across the tokens after the seam — it concentrates right at the start, an attention-sink-style pattern, heavily front-loaded and fading fast. So instead of patching everything, they recompute a small window right after each seam, eight tokens by default, and let the rest ride on the composed state untouched. 20 00:08:29,325 --> 00:08:51,175 [Hal Turing] I'll push back on that one, Ada. A fixed window of eight feels like it was tuned to whatever they happened to benchmark, not something I'd trust blind on a workload nobody's tested. If the deviation decays a little slower on some other segment distribution, you'd silently under-correct and nobody would know until quality quietly drops. 21 00:08:51,175 --> 00:09:12,950 [Dr. Ada Shannon] I actually disagree with you there, Hal. Eight isn't a number they pulled out of thin air — it lines up with the attention-sink literature, where the first handful of tokens absorb disproportionate attention mass regardless of model or task. That's not benchmark-specific, that's a structural property of early-sequence attention allocation. And they're not claiming zero risk past the window, they're claiming the residual past token eight is small enough to eat rather than correct. 22 00:09:12,950 --> 00:09:29,100 [Hal Turing] Fair, and I'll grant the sink pattern is well established elsewhere. I still want it stress-tested outside their four models before I fully trust it, but I take the point that it's grounded in something real, not a knob twisted until the eval looked good. 23 00:09:29,100 --> 00:10:33,225 [Dr. Ada Shannon] Agreed, fair place to land. Last piece is making all of this fast — segment parallelism. Cold segments get scattered to separate workers that prefill them independently, then their end-states get combined back through the same composition law. The scheduling problem is classic load balancing — segments vary a lot in length, so they assign the longest ones first, LPT, so no single worker becomes the straggler. Put it together and here's what it buys: 3.25x average reduction in time-to-first-token, 1.66x more sustainable queries per second at a one-second SLO, and a 5.7x cold-prefill speedup at eight workers. It's not free — average quality lands 1.71 points behind Full Recompute, their from-scratch baseline. And one more number worth sitting with: isolating just the linear-state drift, with full-attention layers disabled so only the recurrent path is measured, deep-layer fidelity comes out to 8.92% relative L2 error and about 5.11 degrees of angular deviation on Qwen3.5-35B-A3B, and a similar 8.69% and 4.98 degrees on Ring-flash. 24 00:10:33,225 --> 00:10:49,500 [Hal Turing] That eight-workers number is going to stick with me — that's the kind of speedup that actually changes whether you'd deploy this. Alright, hold that 1.71-point gap in your head, Ada, because next up I want to know exactly where it's coming from. 25 00:10:49,500 --> 00:11:30,350 [Hal Turing] Alright, Ada, the 1.71-point gap. Table 4 gives you a number for where it comes from — 8.92 percent relative L2, 5.11 degrees at the deepest linear layer on Qwen3.5-35B-A3B, 8.69 percent and 4.98 degrees on Ring-flash. But I went back and reread that methodology note twice because something felt off — they disabled every full-attention layer to get that number. So the 1.71 points listeners saw in the accuracy table is the whole deployed stack, seam windows and all, but the drift number you're about to explain it with isn't measured on that same system. 26 00:11:30,350 --> 00:12:12,475 [Dr. Ada Shannon] That's exactly right, and it's worth being precise about why they did it that way. They wanted to isolate linear-state drift specifically — leave full-attention layers active and seam-window recomputation partially masks whatever error accumulated underneath, so you can't tell how much came from where. Disabling full attention gives a clean read on one failure mode alone. But that's the gap: nowhere do they run it the other direction — full stack, seams on, then decompose the 1.71-point loss into linear-state drift versus residual seam error versus interaction between the two. The 8.92 percent is a best-case, single-mechanism number. It proves linear composition doesn't blow up in isolation. It doesn't tell you what happens once both approximations compound for real. 27 00:12:12,475 --> 00:12:36,475 [Hal Turing] Oh — wait, hold on, sorry to jump in — but doesn't that cut both ways? If seam windows already fix the worst attention-sink deviation right at the segment join, and that's also where linear-state drift is freshest and smallest since less time has passed since the zero-start state, those two errors could be correlated enough to partially cancel rather than stack. We genuinely don't know which. 28 00:12:36,475 --> 00:13:16,400 [Dr. Ada Shannon] Right, and that's the honest answer — we don't, because they never ran that decomposition. It could go either way, and citing 8.92 percent as the drift number is doing more work than it's earned. The exact same shape of gap shows up with family coverage, actually. Table 1's algebra and Table 3's storage numbers treat scalar, diagonal, and dense uniformly — that's the unified claim. But look at what actually ran end to end: Ring-mini and Ring-flash are scalar decay, Qwen3.5-35B and 122B use KDA, which is dense. Nobody ran a GLA-style diagonal model through the full pipeline. Diagonal is proven on paper, not in production. 29 00:13:16,400 --> 00:13:40,000 [Hal Turing] I actually disagree that's a real gap, Ada. Diagonal sits structurally between scalar and dense in Table 1 — it's scalar generalized to a data-dependent gate, and dense with the outer-product erasure term zeroed out. If composition holds exactly on both sides of it, I'd bet real money it holds for diagonal too. That's not hand-waving, that's the shape of the proof. 30 00:13:40,000 --> 00:14:17,550 [Dr. Ada Shannon] No no no, that's not how I'd bet it, Hal. Nested algebra doesn't mean nested failure modes — GLA's gating vector lives per-channel, so composition error could concentrate unevenly across dimensions in a way neither the scalar nor the low-rank dense case would expose. Gated Delta Networks — Yang, Kautz, and Hatamizadeh out of NVIDIA, 2025 — is actually why the dense case looks so clean here: KDA stacks a diagonal gate on top of the delta rule, so Qwen3.5 already exercises a hybrid of dense and diagonal behavior without anyone calling it that. I'll grant your instinct is reasonable. I wouldn't ship it as proven. 31 00:14:17,550 --> 00:14:57,075 [Hal Turing] Fair, we can agree to disagree on how much that risk actually bites in practice. Zooming out, though — if you're running RAG or agentic serving on Qwen3.5, Ring-2.5, or Kimi-Linear today, what does this actually buy you? It's built directly on CacheBlend's splice-and-correction primitives — Yao, Li, Liu and colleagues out of the University of Chicago, EuroSys 2025 — extended into a setting CacheBlend was never designed for. If you're already running CacheBlend-style PIC on a full-attention model and migrating to a hybrid architecture for the compute savings, this is the first path that doesn't force you to give PIC up entirely. 32 00:14:57,075 --> 00:15:54,550 [Dr. Ada Shannon] And the seam-window insight has a direct lineage too — it extends the attention-sink locality finding from EPIC, Junhao Hu's group at Peking University, 2025, which showed that same beginning-of-segment concentration in pure full-attention stacks. Junhao's a co-author here too, continuing that line directly. What's still open: the decomposed error attribution we just argued about, validating GLA end to end, and testing beyond RAG and summarization — the intro motivates this with agentic serving, but nothing in the eval touches a multi-turn agent trace. There's also a gap the paper never touches — the public pool shares cached segments across tenants with no isolation discussion. If two customers' RAG pipelines embed overlapping documents, could the shared cache become a timing side channel revealing what someone else queried? Full-attention KV-cache sharing has already drawn that scrutiny; this coarser cache hasn't been checked at all. 33 00:15:54,550 --> 00:16:32,200 [Hal Turing] Good note to end on. So: Hypic gives you constant-time state composition for the linear layers through that cached transition operator, a cheap seam window to patch the full-attention layers, and segment parallelism that turns cold requests from a tail-latency liability into something you can just throw more workers at. Real speedups, an honestly reported quality cost, and a couple of claims — the drift decomposition, the diagonal family — staked out a bit further than the evidence currently reaches. That's Hypic. Thanks for listening, everyone — we'll catch you next time.