arXiv:2604.02473 GPU fabrics Address translation MoE / All-to-All

Small Collectives, Big TLB Cost

Reverse Address Translation in Scale-Up GPU Pods — when a remote request lands on a GPU carrying an address that means nothing locally, someone has to translate it, with no warning and no control over the pattern.

Amel Fatima, Tuan Ta, Bradford M. Beckmann — University of Virginia & AMD Research · arXiv April 2026

Mailing a letter vs. reaching into a drawer

RDMA hands data to a NIC, which pushes it over the network to a remote NIC. Scale-up fabrics like NVLink and UALink cut the NIC out — but direct access means the destination GPU receives a raw fabric address it has never seen before.

idle path active hop translation required

The NPA → SPA hierarchy (hover each stage)

The source GPU's MMU turns a virtual address into a Network Physical Address — a shipping label, not a street address. Only the destination's Link MMU can turn it back into something memory understands.

All-to-All dispatch & gather — twice per layer

Every GPU sends a distinct chunk to every other GPU. In Mixture-of-Experts, dispatch routes tokens out to experts, gather brings results back — each crossing the translation step. Click a GPU to trace its edges.

dispatch gather RAT crossing

Cumulative RAT crossings by layer depth

A 6-GPU illustrative pod: each layer forces 2 × (N−1) reverse translations per GPU. Stack that across dozens of layers and the "niche hardware detail" stops being niche.

Execution-time degradation vs. an ideal, zero-overhead baseline

Small collectives suffer most because nearly every request walks a cold page table. Hover a cell — pod size on rows, collective size on columns.

~1.0x (low) ~1.2x ~1.4x (high)

Reverse Address Translation as % of round-trip latency

At 1MB, roughly 30% of per-request round-trip time is spent purely on RAT. Once a collective is big enough to reuse warmed entries, that cost amortizes toward zero.

The L1-MSHR trap

Over 90% of inter-node requests hit the L1-MSHR — sounds great, but a hit can still stall behind a pending walk underneath. Toggle collective size to see what's really happening below that headline number.

Cold-miss rate over a 256MB collective

One spike of cold misses at the very start, then flat — each GPU streams sequentially through one page per source and rarely revisits it.

L2 Link TLB sweep — 16 to 32,768 entries (32-GPU pod, 16MB collective)

A thousand-x sweep lands on the same number intuition predicts: 32 entries, matched to GPU count, already saturates performance. Bigger buys nothing.

Working set = 1 page per source GPU

Pick a pod size — the L2's required working set scales with GPU count, not collective size.

What's tested vs. extrapolated

The mechanism looks architectural; the exact numbers are baseline- and scale-dependent. Green = directly tested, orange = plausible but untested, red = not addressed.

References