Beluga is a memory-topology paper: the transformer stays the same, but the bottleneck moves into where KV cache lives, how it is fetched, and how much software ceremony sits between a GPU and reusable context state.
Transcript scan found one explicit arXiv-style identifier. The rest of the companion focuses on the system envelope: local HBM, host DRAM, remote pooled memory, and the reuse conditions where Beluga’s wins are likely to be real.
Toggle the retrieval path. The visual keeps the same serving task and changes only the memory access model: staged RDMA versus a CXL shared pool with flatter data movement.
The heatmap shows which KV segments get reused across tenants and requests. Hover cells to inspect pressure, then walk the access path step by step.
These mock-but-structured curves illustrate the same caution voiced in the episode: cache-hit-heavy workloads can look spectacular, while cold-start writes stay bounded by slower remote tiers.
Best case here means large prefix overlap, many repeated retrievals, and enough queue depth to amortize setup overhead. The shape matters more than the exact number.
Pooling, compression, scheduling, and placement attack the same bottleneck from different angles. The map below positions Beluga against nearby systems ideas discussed in the episode.
Beluga looks strongest when the operator can control rack topology, expects heavy reuse, and prefers simpler load/store semantics over a more protocol-heavy RDMA control plane.
Compact pointers to the paper, adjacent CXL work, and related KV-cache system directions.