System Architecture
Every request splits across a Prefill Instance and a Decode Instance. The question the paper asks is: what sits between them, and how does the decode side get the top-k entries it actually needs? Toggle between the RDMA baseline (Mooncake/LMCache-style) and SAC's CXL design.
Prefill → KV Pool → Decode
Same three-stage pipeline, two very different middle layers.
DeepSeek Sparse Attention: What Actually Gets Read
The Lightning Indexer scores every historical KV entry cheaply, then keeps only the top-2048 per layer — roughly 1.6% of a 128K-token context. This grid is a down-sampled illustration of that ratio (576 sampled positions, 9 selected) so the sparsity is visible at a glance.
Indexer Score Heatmap
Each cell is one KV entry. Hover a cell for its indexer score. Toggle to see what "attention" actually means under each regime.
Why RDMA Can't Cheaply Do Scattered Top-K Lookups
RDMA is a message-based protocol built for moving one big contiguous block. CXL gives the GPU hardware-managed load/store at cache-line granularity — it can just read memory like local DRAM.
Measured Latency vs Local DRAM
Range bars: reported min–max multiplier over a local DRAM access, per the paper's microbenchmarks.
Protocol Path, Step by Step
Toggle the transport to see what has to happen before a byte is usable.
End-to-End Numbers on DeepSeek-V3.2
Reported gains of SAC (CXL) over the RDMA baseline. Throughput is higher-is-better; TTFT and TBT are lower-is-better — watch which direction each bar wants to go.
Baseline vs SAC, by Metric
Bars normalized so RDMA baseline = 100 on every metric.
Interconnect Cost, per 64GB/s
Table 3 in the paper. Excludes the terabyte-scale local DRAM over-provisioning RDMA disaggregation typically requires.
Reading the Fine Print
The headline numbers hold up at the microbenchmark level. The disaggregation story is a different question — here's what the test rig actually was versus what gets marketed.
Marketed Topology vs Tested Topology
Figure 7 markets "up to 8 servers" sharing a CXL switch. Every number in Sections 5.1–5.4 comes from one physical box.
Tail Latency vs Concurrency (Appendix D.3)
Mean grows modestly; p99 grows faster, and the gap is wider for CXL than local DRAM at equal concurrency. Max tested concurrency was 192 — the shaded zone beyond it is unmeasured.
References
| 1 | SAC: Disaggregated KV Cache System for Sparse Attention
LLMs with CXL — Ma, Ma, Li, Zha, Shang, Hu, Liu, Yang, Ma, Luo.
arxiv.org/abs/2606.19746 |
2026 |
| 2 | Mooncake: Trading More Storage for Less Computation
— A KVCache-centric Architecture for Serving LLM Chatbot —
Qin, Li, He, et al.
Google Scholar |
2025 |
| 3 | Beluga: A CXL-based Memory Architecture for Scalable
and Efficient LLM KVCache Management — Yang, Hu, Li, Li, et al.
Google Scholar |
2026 |
| 4 | TraCT: Disaggregated LLM Serving with CXL Shared
Memory KV Cache at Rack-Scale — Yoon, Min, Kim, Noh, Kim.
Google Scholar |
2025 |
| 5 | HiSparse: High-Efficiency Sparse Attention Inference
in SGLang — Xie, Huang, Huang.
Google Scholar |
2026 |
| 6 | SnapKV: LLM Knows What You Are Looking For Before
Generation — Li, Huang, Yang, et al.
Google Scholar |
2024 |
| 7 | H2O: Heavy-Hitter Oracle for Efficient Generative
Inference of Large Language Models — Zhang, Sheng, Zhou, et al.
Google Scholar |
2023 |