1 00:00:01,000 --> 00:00:32,225 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called SAC: Disaggregated KV Cache System for Sparse Attention LLMs with CXL. That's Ruiyang Ma et al. — ten authors total — out of Peking University, Alibaba Cloud, and Renmin University of China. It went up on arXiv on June 18th, 2026. Ada, you flagged this one to me pretty fast. 2 00:00:32,225 --> 00:01:04,150 [Dr. Ada Shannon] Yeah, most systems papers in this space give you a clever engineering trick and call it a day. This one actually builds a case from first principles — they measure the raw memory-access latency of the hardware before they design anything, and the design falls out of that measurement almost inevitably. That's rare. Usually I'm skeptical of papers that claim a hardware swap alone buys you an order-of-magnitude win, but here the groundwork is solid enough that I bought it before I even got to their end-to-end numbers. 3 00:01:04,150 --> 00:01:17,376 [Hal Turing] Okay, so let's set the stage, because I think a lot of listeners know 'attention is compute-bound' as a kind of default assumption, and this paper is built on that no longer being true. What changed? 4 00:01:17,376 --> 00:01:57,701 [Dr. Ada Shannon] Right, so historically, LLM serving was compute-bound — you were waiting on the GPU to crunch matrix multiplies. But as context lengths blew up, the bottleneck shifted to memory: specifically, the KV cache, which stores the key-value pairs from every prior token so the model doesn't recompute attention from scratch each step. At long contexts that cache can hit tens of gigabytes per request, way past what fits in GPU HBM. So the industry standard became disaggregated KV cache — you park the cache in a separate memory pool, reachable over a fabric, and systems like Mooncake and LMCache prefetch the entire thing over RDMA before decoding starts. 5 00:01:57,701 --> 00:02:17,251 [Hal Turing] And that works fine as long as the model actually needs the whole cache. But that's exactly the assumption sparse attention breaks, right? If the model only ever looks at a fraction of the context, hauling the entire thing across the wire every single time sounds like a lot of wasted motion to me. 6 00:02:17,251 --> 00:02:47,701 [Dr. Ada Shannon] Exactly — for dense attention, every token attends to every prior token, so fetching the full prefix is the correct thing to do. But sparse attention models like DeepSeek-V3.2 use something called DeepSeek Sparse Attention, or DSA, where a lightweight 'Lightning Indexer' scores every historical KV entry and only the top-k — 2048 of them — actually get used per layer. That drops compute from quadratic to near-linear. Great for FLOPs. Terrible fit for an RDMA pipeline that insists on hauling the entire prefix cache to local memory regardless. 7 00:02:47,701 --> 00:02:59,601 [Hal Turing] Oh wait wait wait — so you're telling me the model only touches a tiny sliver of what got fetched? That feels like the whole point of sparse attention just got undone by the plumbing. 8 00:02:59,601 --> 00:03:25,951 [Dr. Ada Shannon] That's the paper's whole motivation in one sentence. They name two concrete failure modes. P1 is the transmission bottleneck — moving dozens of gigabytes over the network per request chokes bandwidth and queues requests, hurting time-to-first-token. P2 is local memory wasting — you still have to hold the full cache locally even though, at 128K context, only a sliver of it ever gets touched during decoding. So you're burning both network and memory for data you'll mostly never read. 9 00:03:25,951 --> 00:03:39,726 [Hal Turing] So why not just fetch the top-k entries on demand instead of the whole cache then? Seems like the obvious fix once you put it that way — skip the full haul, grab only what the model's actually going to attend to. 10 00:03:39,726 --> 00:04:15,551 [Dr. Ada Shannon] It is obvious, and it's also infeasible over RDMA, which is the interesting twist. RDMA is a message-based protocol — send and receive, with queue pairs, pinning, and context switching stacked on top. Great for moving one big contiguous block, terrible for dozens of tiny scattered top-k lookups per layer at real-time latency. That's where CXL, Compute Express Link, comes in. It's a PCIe-based interconnect that gives you hardware-managed load and store semantics at cache-line granularity — the GPU can just read memory like it's local DRAM, no message protocol overhead at all. 11 00:04:15,551 --> 00:04:25,301 [Hal Turing] And the headline numbers back that up, right? I saw a 2.1x figure in my notes before we started recording — what's the full picture on that? 12 00:04:25,301 --> 00:05:20,176 [Dr. Ada Shannon] On DeepSeek-V3.2, their system SAC gets 2.1x higher throughput, 9.7x lower time-to-first-token, and 1.8x lower time-between-tokens compared to the RDMA baseline. TTFT and TBT are the two latency metrics that matter most for anyone actually serving these models, so those aren't small wins — that's the difference between a snappy product and a frustrating one. And it's deliberate — the system is built in three pieces that hand off to each other. There's a Prefill Instance, a Decode Instance built on top of something called HiSparse, and the CXL-based disaggregated KV cache system that ties the two together. The prefill side is honestly the boring part — full KV blocks get computed and written into the CXL pool, and that's a solved problem from prior CXL work. All the hard engineering sits on the decode side, because that's where you're chasing a moving target of top-k indices, per layer, in real time, not just writing a static block once. 13 00:05:20,176 --> 00:05:53,226 [Hal Turing] So that's the system end to end. Now let's poke at how it was actually tested, because Appendix A.1 has a detail I can't let slide. The RDMA baseline isn't real cross-node RDMA — it's eight ConnectX-7 NICs looped back on themselves, KV cache sitting in local DRAM the whole time. The paper calls that 'conservative, best-case' for RDMA, meaning their real-world gains are supposedly bigger than reported. But the CXL side is also best-case — one switch, one hop. Is this actually a fair fight? 14 00:05:53,226 --> 00:06:32,401 [Dr. Ada Shannon] Fair question, and the answer cuts both ways. Loopback RDMA does dodge what actually kills RDMA in production — multi-hop switch traversal, cable propagation, real contention from other tenants on the fabric. So yes, a genuine multi-hop deployment would look worse than what's reported, which favors SAC's argument. But the CXL topology has the mirror problem — single switch, zero contention from other hosts pulling on the pool. Neither baseline gets stress-tested for what a real multi-tenant deployment looks like. The relative gap, 2.1x and 9.7x, probably survives. Whether the absolute numbers do once you add real topology depth on both sides — that's untested. 15 00:06:32,401 --> 00:06:57,476 [Hal Turing] Wait — hold on, that same question hits even harder somewhere else. Figure 7 markets 'up to 8 servers' sharing one CXL switch — that's the whole disaggregation pitch. But Appendix A.1's actual rig is one physical server, eight H20 GPUs, one switch, one pool. That's not eight servers contending for shared memory. That's one box with an external memory expansion chassis. 16 00:06:57,476 --> 00:08:01,102 [Dr. Ada Shannon] Right, and I checked — every number in Sections 5.1 through 5.4 comes from that single 8-GPU box. 'Up to 8 servers' describes what the switch chip supports, 256 PCIe lanes, not what was run. So 'disaggregation' here is entirely intra-server. Same pattern with model coverage — every result is DeepSeek-V3.2 in AWQ 4-bit. GLM-5.1 and DeepSeek-V4 are explicitly future work, so 'superior infrastructure for sparse attention models' as a class is one model backing a claim about a family. Worth grounding against where this field came from, too. Mooncake — Ruoyu Qin and coauthors, Moonshot AI, FAST 2025 — is the production RDMA system everyone measures against, and it's also the source of SAC's own 100-to-1 agentic input-output ratio figure. Beluga — Xinjun Yang, Qingda Hu, Junru Li and coauthors, 2026, sharing authors with this paper — already proved CXL beats RDMA for dense block transfers. SAC's Appendix B.1.3 admits it: this is 'transfer to on-demand fetching,' not a new insight about CXL itself. 17 00:08:01,102 --> 00:08:28,902 [Hal Turing] There's one more thing buried in Appendix D.3 — tail latency. Mean TBT and TTFT grow modestly with concurrency, but p99 grows faster, and the mean-to-p99 gap is wider for CXL than local DRAM at equal concurrency. They chalk it up to arbitration overhead in the fabric. Max concurrency tested was 192. We have no idea if that gap stays linear or blows up at production concurrency in the thousands. 18 00:08:28,902 --> 00:09:03,177 [Dr. Ada Shannon] And that's exactly the kind of thing that bites harder at scale than in a single-server test. Economically though, it's hard to argue with — Table 3 puts RDMA interconnect at 800 dollars per 64 gigabytes a second, versus $218.75 for the CXL equivalent, plus you're not over-provisioning terabyte-scale local DRAM per node. If you're running sparse-attention models today and you're RDMA-bound on KV transmission, this is worth evaluating — just evaluate it as a single-node memory-tier swap, not a proven multi-host pooling fabric. 19 00:09:03,177 --> 00:09:42,902 [Hal Turing] So here's where I land. The microbenchmark case is genuinely convincing — CXL at 1.04 to 1.64x of local DRAM latency versus RDMA's 4 to almost 20x isn't close, and the single-node numbers back it up: 2.1x throughput, 9.7x TTFT, 1.8x TBT. But the multi-host disaggregation story and the generality to other sparse models are ahead of what was actually measured. Worth watching for a follow-up that puts real independent servers on that switch. That's SAC — CXL for sparse attention KV caches. Thanks for listening, everyone.