Interactive Visualization arXiv: 2511.20172 Viz URL

Beluga: CXL Memory Pooling for LLM KV Cache

Beluga is a memory-topology paper: the transformer stays the same, but the bottleneck moves into where KV cache lives, how it is fetched, and how much software ceremony sits between a GPU and reusable context state.

Core Claim
CXL Pool
Shared load/store-style KV access instead of explicit RDMA fetch choreography.
Reported Best TTFT
-89.6%
Largest gains appear in cache-hit-heavy retrieval regimes.
Reported Best Throughput
7.35×
Compared against the paper’s RDMA-style baseline path.

Extracted arXiv IDs

2511.20172

Transcript scan found one explicit arXiv-style identifier. The rest of the companion focuses on the system envelope: local HBM, host DRAM, remote pooled memory, and the reuse conditions where Beluga’s wins are likely to be real.

From “networked fetches” to “mapped remote memory”

Toggle the retrieval path. The visual keeps the same serving task and changes only the memory access model: staged RDMA versus a CXL shared pool with flatter data movement.

GPU / local compute Shared pool / fabric Explicit staging or copy pressure

Latency Ladder

KV locality is not binary

The heatmap shows which KV segments get reused across tenants and requests. Hover cells to inspect pressure, then walk the access path step by step.

Prefix-Reuse Heatmap

Read / Write Pipeline

When the headline speeds are plausible

These mock-but-structured curves illustrate the same caution voiced in the episode: cache-hit-heavy workloads can look spectacular, while cold-start writes stay bounded by slower remote tiers.

Reuse Regime Explorer

Best case here means large prefix overlap, many repeated retrievals, and enough queue depth to amortize setup overhead. The shape matters more than the exact number.

Beluga sits in a larger design space

Pooling, compression, scheduling, and placement attack the same bottleneck from different angles. The map below positions Beluga against nearby systems ideas discussed in the episode.

Interpretation

Beluga looks strongest when the operator can control rack topology, expects heavy reuse, and prefers simpler load/store semantics over a more protocol-heavy RDMA control plane.

References

Compact pointers to the paper, adjacent CXL work, and related KV-cache system directions.