AI Post Transformers · Episode Companion

SAC: Making Sparse Attention KV Caches Actually Sparse with CXL

Ruiyang Ma, Teng Ma, Junru Li, Hantian Zha, Xuchun Shang, Qingda Hu, Zheng Liu, Xinjun Yang, Tao Ma, Guojie Luo · Peking University, Alibaba Cloud, Renmin University of China · June 18, 2026
arXiv:2606.19746 Disaggregated KV Cache CXL vs RDMA DeepSeek-V3.2 / DSA

System Architecture

Every request splits across a Prefill Instance and a Decode Instance. The question the paper asks is: what sits between them, and how does the decode side get the top-k entries it actually needs? Toggle between the RDMA baseline (Mooncake/LMCache-style) and SAC's CXL design.

Prefill → KV Pool → Decode

Same three-stage pipeline, two very different middle layers.

Dense attention needs the full prefix, so hauling everything over RDMA is the correct behavior for dense models. Sparse attention breaks that assumption — the decode side only ever reads a top-k slice per layer.

DeepSeek Sparse Attention: What Actually Gets Read

The Lightning Indexer scores every historical KV entry cheaply, then keeps only the top-2048 per layer — roughly 1.6% of a 128K-token context. This grid is a down-sampled illustration of that ratio (576 sampled positions, 9 selected) so the sparsity is visible at a glance.

Indexer Score Heatmap

Each cell is one KV entry. Hover a cell for its indexer score. Toggle to see what "attention" actually means under each regime.

Low indexer score Mid score High score Scored but discarded (sparse mode only)
Under RDMA disaggregation, the entire prefix cache still gets hauled to local memory regardless of this heatmap — the sparsity the model computes never reaches the network layer. That's P1 (wasted bandwidth) and P2 (wasted local memory) in the paper's framing.

Why RDMA Can't Cheaply Do Scattered Top-K Lookups

RDMA is a message-based protocol built for moving one big contiguous block. CXL gives the GPU hardware-managed load/store at cache-line granularity — it can just read memory like local DRAM.

Measured Latency vs Local DRAM

Range bars: reported min–max multiplier over a local DRAM access, per the paper's microbenchmarks.

Protocol Path, Step by Step

Toggle the transport to see what has to happen before a byte is usable.

The gap isn't marginal: CXL sits at roughly 1.0–1.6× local DRAM latency; RDMA sits at 4–20×. That's the physical basis for everything downstream in the Results tab.

End-to-End Numbers on DeepSeek-V3.2

Reported gains of SAC (CXL) over the RDMA baseline. Throughput is higher-is-better; TTFT and TBT are lower-is-better — watch which direction each bar wants to go.

Baseline vs SAC, by Metric

Bars normalized so RDMA baseline = 100 on every metric.

RDMA Baseline SAC (CXL)

Interconnect Cost, per 64GB/s

Table 3 in the paper. Excludes the terabyte-scale local DRAM over-provisioning RDMA disaggregation typically requires.

Reading the Fine Print

The headline numbers hold up at the microbenchmark level. The disaggregation story is a different question — here's what the test rig actually was versus what gets marketed.

Marketed Topology vs Tested Topology

Figure 7 markets "up to 8 servers" sharing a CXL switch. Every number in Sections 5.1–5.4 comes from one physical box.

Tail Latency vs Concurrency (Appendix D.3)

Mean grows modestly; p99 grows faster, and the gap is wider for CXL than local DRAM at equal concurrency. Max tested concurrency was 192 — the shaded zone beyond it is unmeasured.

p50 latency p99 latency Untested → production-scale concurrency
Economics still favor CXL even after the scrutiny: $218.75 vs $800 per 64GB/s, plus no need to over-provision terabyte-scale local DRAM per node. Evaluate it as a single-node memory-tier swap, not yet a proven multi-host pooling fabric.

References

1 SAC: Disaggregated KV Cache System for Sparse Attention LLMs with CXL — Ma, Ma, Li, Zha, Shang, Hu, Liu, Yang, Ma, Luo.
arxiv.org/abs/2606.19746
2026
2 Mooncake: Trading More Storage for Less Computation — A KVCache-centric Architecture for Serving LLM Chatbot — Qin, Li, He, et al.
Google Scholar
2025
3 Beluga: A CXL-based Memory Architecture for Scalable and Efficient LLM KVCache Management — Yang, Hu, Li, Li, et al.
Google Scholar
2026
4 TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale — Yoon, Min, Kim, Noh, Kim.
Google Scholar
2025
5 HiSparse: High-Efficiency Sparse Attention Inference in SGLang — Xie, Huang, Huang.
Google Scholar
2026
6 SnapKV: LLM Knows What You Are Looking For Before Generation — Li, Huang, Yang, et al.
Google Scholar
2024
7 H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhang, Sheng, Zhou, et al.
Google Scholar
2023