Mailing a letter vs. reaching into a drawer
RDMA hands data to a NIC, which pushes it over the network to a remote NIC. Scale-up fabrics like NVLink and UALink cut the NIC out — but direct access means the destination GPU receives a raw fabric address it has never seen before.
The NPA → SPA hierarchy (hover each stage)
The source GPU's MMU turns a virtual address into a Network Physical Address — a shipping label, not a street address. Only the destination's Link MMU can turn it back into something memory understands.
All-to-All dispatch & gather — twice per layer
Every GPU sends a distinct chunk to every other GPU. In Mixture-of-Experts, dispatch routes tokens out to experts, gather brings results back — each crossing the translation step. Click a GPU to trace its edges.
Cumulative RAT crossings by layer depth
A 6-GPU illustrative pod: each layer forces 2 × (N−1) reverse translations per GPU. Stack that across dozens of layers and the "niche hardware detail" stops being niche.
Execution-time degradation vs. an ideal, zero-overhead baseline
Small collectives suffer most because nearly every request walks a cold page table. Hover a cell — pod size on rows, collective size on columns.
Reverse Address Translation as % of round-trip latency
At 1MB, roughly 30% of per-request round-trip time is spent purely on RAT. Once a collective is big enough to reuse warmed entries, that cost amortizes toward zero.
The L1-MSHR trap
Over 90% of inter-node requests hit the L1-MSHR — sounds great, but a hit can still stall behind a pending walk underneath. Toggle collective size to see what's really happening below that headline number.
Cold-miss rate over a 256MB collective
One spike of cold misses at the very start, then flat — each GPU streams sequentially through one page per source and rarely revisits it.
L2 Link TLB sweep — 16 to 32,768 entries (32-GPU pod, 16MB collective)
A thousand-x sweep lands on the same number intuition predicts: 32 entries, matched to GPU count, already saturates performance. Bigger buys nothing.
Working set = 1 page per source GPU
Pick a pod size — the L2's required working set scales with GPU count, not collective size.
What's tested vs. extrapolated
The mechanism looks architectural; the exact numbers are baseline- and scale-dependent. Green = directly tested, orange = plausible but untested, red = not addressed.