A visual-first walk through GPU-native memory expansion: local HBM, host-managed UVM, CXL DRAM, and SSD-backed tiers. The paper’s strongest claim is not “remote memory becomes local,” but that GPU-side CXL controllers, multiple root ports, speculative reads, and deterministic stores can make far tiers less painful than host-mediated migration.
The hierarchy is the point. Capacity rises as you move outward, while bandwidth and immediacy fall. Hover the cells and lanes: the page is showing why “memory-like” is useful without pretending all tiers behave the same.
The left chart compresses a systems argument into geometry: HBM is narrow and high on the “instant” axis, SSD-backed expansion is huge but delayed, and CXL DRAM tries to open a practical middle zone.
2506.15601 was the only explicit arXiv-style ID in the transcript.
Step through a miss leaving the GPU. The point is not magic prediction. The point is inserting speculative read, queue-aware throttling, and deterministic store where the GPU can hide some of the pain before host software turns every miss into a larger event.
These charts use realistic mock data shaped by the episode’s framing: strong gains on irregular out-of-core kernels, milder gains against newer paths, and a visible gap between classic heterogeneous benchmarks and transformer-like access behavior.
The episode’s unresolved question becomes a heatmap. The paper likely helps some overflow patterns. It is much less clear that it solves the nastiest modern workloads: long-context serving, bursty KV growth, activation rematerialization, and multi-port contention at scale.