1 00:00:01,000 --> 00:00:31,125 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper is 'Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods' — Amel Fatima et al., three authors total: Amel Fatima, Tuan Ta, and Bradford M. Beckmann, out of the University of Virginia and AMD Research. It went up on arXiv April 2nd, 2026. Ada, on paper this sounds like pure hardware plumbing. Sell me on it. 2 00:00:31,125 --> 00:00:53,200 [Dr. Ada Shannon] Happily. Everyone obsesses over compute — bigger GPUs, faster links — and nobody asks what happens the instant a remote request lands on the receiving GPU carrying an address that means nothing locally. This is the first paper to actually measure that gap instead of hand-waving about it in a spec footnote. What sold me is they built the translation hierarchy from first principles and reasoned through why it behaves the way it does, rather than just running a benchmark and calling it a day. 3 00:00:53,200 --> 00:01:07,924 [Hal Turing] Okay, ground me first, because I want the non-hardware folks with us too. Why do we even need something like NVLink or UALink, instead of just going over the network the way we always have with RDMA and InfiniBand? 4 00:01:07,924 --> 00:01:45,000 [Dr. Ada Shannon] RDMA is mediated — your GPU hands data to a NIC, the NIC pushes it over InfiniBand or Ethernet, and the remote NIC hands it to the remote GPU. Extra hop, extra latency, every time. Scale-up fabrics cut the NIC out entirely. NVLink and the UALink standard — ratified by an industry consortium in 2025 — give GPUs direct load, store, and atomic access straight into each other's memory, across pods of up to 1,024 accelerators. It's the difference between mailing a letter and reaching straight into someone's desk drawer. And that's exactly what creates the problem this paper is about. 5 00:01:45,000 --> 00:01:52,674 [Hal Turing] Reaching into the drawer, sure — but if it's direct access, why do you need translation at all on the receiving end? 6 00:01:52,674 --> 00:02:12,050 [Dr. Ada Shannon] Because the address that crosses the fabric isn't a real address yet. The source GPU's own MMU translates its virtual address into either a System Physical Address if the target is local, or something new if it's remote — a Network Physical Address, an NPA. And an NPA only means something on the network. It's not— 7 00:02:12,050 --> 00:02:21,050 [Hal Turing] Oh wait wait wait — it's not an actual address at the destination at all, is it? It's more like a shipping label than a street address. 8 00:02:21,050 --> 00:02:47,925 [Dr. Ada Shannon] Exactly. So the destination GPU converts that NPA back into its own local SPA before it can touch memory — that conversion is Reverse Address Translation, the mirror image of the virtual-to-physical work your CPU does on every access, except it happens at the target, triggered by someone else's request. Both specs acknowledge they need a Link MMU with a Link TLB hierarchy to do this, but neither says much about how to build it well. That silence is basically the whole motivation for the paper. 9 00:02:47,925 --> 00:02:58,650 [Hal Turing] TLB design isn't exactly new territory though — people have optimized translation for decades. Why is the destination side such a blank page? 10 00:02:58,650 --> 00:03:29,775 [Dr. Ada Shannon] Because that entire decades-long lineage assumes the initiator does the translating. Bhattacharjee and Martonosi's Inter-Core Cooperative TLB work out of Princeton, ASPLOS 2010, or Barr, Cox, and Rixner's SpecTLB out of Rice, 2011 — all of it optimizes the GPU or CPU issuing the access. Reverse Address Translation flips that: the GPU being accessed has to do the work, with zero warning and zero control over the pattern. There was no destination-side translation step at this scale to study before scale-up fabrics existed. 11 00:03:29,775 --> 00:03:40,825 [Hal Turing] I'll be honest, Ada, my gut says this is still a narrow corner of the stack — one more module inside a fabric most engineers will never touch. Worth a whole paper? 12 00:03:40,825 --> 00:03:56,025 [Dr. Ada Shannon] I actually disagree with you there, and pretty strongly. This sits on the critical path of every All-to-All collective in every Mixture-of-Experts model shipping right now. That's not a corner case, that's the workload everyone's racing to scale. 13 00:03:56,025 --> 00:04:06,800 [Hal Turing] Maybe, but that's one piece among dozens adding latency — compute, scheduling, cache movement. I need to see it's meaningfully large before I call it critical. 14 00:04:06,800 --> 00:04:27,074 [Dr. Ada Shannon] Fair, and I'm not pretending I've shown you the receipts yet — that's next. But structurally, an MoE layer does All-to-All dispatch and All-to-All gather, twice per layer, dozens of layers deep, and every one of those crosses this translation step. You don't opt out of it with a well-tuned kernel elsewhere. So — the paper still has to earn the word 'critical,' but the mechanism itself isn't niche. 15 00:04:27,074 --> 00:04:34,675 [Hal Turing] Okay, that I'll take. Walk me through All-to-All itself, for anyone who hasn't touched MoE routing directly. 16 00:04:34,675 --> 00:05:13,000 [Dr. Ada Shannon] Every GPU sends a distinct chunk of data to every other GPU, all at once — one of the core collective patterns alongside AllReduce and AllGather, implemented in the libraries everyone runs in production: NCCL on NVIDIA, RCCL on AMD, oneCCL on Intel. It's central to Mixture-of-Experts, where each layer routes tokens to a handful of specialist sub-networks instead of running every token through every parameter — two All-to-All exchanges per layer, dispatch out, gather back. Do that across dozens of layers on live inference traffic, and the destination side of that exchange starts to matter a lot. 17 00:05:13,000 --> 00:05:23,425 [Hal Turing] So every one of those dispatch-and-gather round trips leans on this translation step working cleanly. Now I want to know how badly it actually bends. 18 00:05:23,425 --> 00:06:06,225 [Dr. Ada Shannon] Methodologically, they extended ASTRA-sim2.0 — Won, Heo, Rashidi, Sridharan, Srinivasan and Krishna out of Georgia Tech with NVIDIA collaborators, ISPASS 2023 — using Omnet++ as the network backend for packet-level modeling of an actual UALink Clos topology. For the workload, they generate All-to-All traffic with MSCCLang, Microsoft's collective communication language from Cowan and colleagues at Microsoft, ASPLOS 2023, specifically the all-pairs direct algorithm — every GPU fires one chunk straight at every other GPU. And they bake in one explicit assumption: zero cache reuse, zero temporal locality. Every request is modeled as missing at every cache level on the way in. 19 00:06:06,225 --> 00:06:28,300 [Hal Turing] Okay — before the number, quick provenance check, because this matters for how much I trust it: you said UALink and NVLink don't publish real Link MMU specs. So where did the actual TLB hierarchy in their baseline come from, and then hit me with the headline damage number, because I want to know if this is a rounding error or a real problem. 20 00:06:28,300 --> 00:07:32,825 [Dr. Ada Shannon] It's borrowed, and a little cheekily. The whole Link MMU config — 32-entry private L1, 512-entry shared L2, page walk caches, a hundred parallel page table walkers — comes wholesale from Trans-FW, a paper by Bingyao Li and colleagues out of the University of Pittsburgh, HPCA 2023. That paper isn't about UALink at all — it's about short-circuiting page walks in ordinary multi-GPU systems by forwarding requests between GPU IOMMUs. Different problem, but it's the closest thing to a real translation hierarchy anyone's published. Across pod sizes from 8 to 64 GPUs, small 1-megabyte collectives see up to 1.4x execution-time degradation against an ideal, zero-overhead baseline. Sixteen-megabyte collectives only take about 1.1x. The reason is cold-versus-warm caches: at 1 megabyte, roughly 30% of total round-trip latency per request is spent purely on Reverse Address Translation, because nearly every request walks a cold page table. Once a collective is big enough to reuse those same entries across many requests, that cost gets amortized down to almost nothing. 21 00:07:32,825 --> 00:07:49,350 [Hal Turing] That 30% figure is nuts for something buried this deep in the stack most engineers never think about. So walk me through where in the hierarchy that time is actually going — L1, the shared L2, or all the way down at the page walker itself? 22 00:07:49,350 --> 00:08:37,100 [Dr. Ada Shannon] Here's the trap — the L1 number alone will fool you. Over 90% of inter-node requests hit the L1-MSHR, which sounds great, but that stat alone says nothing about latency, because a request can hit the MSHR and still stall behind a pending walk further down. Break down those L1-MSHR hits and for 1-megabyte collectives, L2-TLB misses and hit-under-miss dominate — cold walks stacking up at the bottom. Push past a couple megabytes and it flips hard toward L1-TLB hits instead, meaning the early requests warmed the hierarchy for everyone after them. And the warming has a specific shape too — trace a 256-megabyte collective and you see one big spike of cold misses at the very start, then it flattens, because each GPU streams sequentially through one page and essentially never revisits it once it moves on. So at any given instant, the destination is only ever tracking— 23 00:08:37,100 --> 00:08:55,000 [Hal Turing] Hold on, that's actually the elegant part, isn't it — the destination only ever has to track one live page per source GPU. So the entire working set the L2 needs to cover is bounded by how many GPUs are talking to it, not by how big the collective is. 24 00:08:55,000 --> 00:09:26,500 [Dr. Ada Shannon] Exactly — one page per GPU, full stop, whether the collective is 1 megabyte or 4 gigabytes. And they tested that directly: on a 32-GPU pod, they swept the L2 Link TLB from 16 entries up to 32,768, a thousand-x range, running a 16-megabyte collective. A 32-entry L2, matched exactly to the GPU count, already saturates performance. Bigger buys you nothing — a three-order-of-magnitude sweep landing on the exact same number your intuition just derived from first principles. 25 00:09:26,500 --> 00:09:42,250 [Hal Turing] So then the real takeaway is: don't bother building a big L2 at all — a small, fixed TLB is all you'll ever need no matter how the pod scales. The CPU-world instinct that bigger cache is always better just doesn't apply here. 26 00:09:42,250 --> 00:10:09,424 [Dr. Ada Shannon] I actually disagree with how you're framing that, Hal. It's not 'a small TLB' — it's a TLB sized to your GPU count, and those are different design philosophies. They showed 32 entries is enough for 32 GPUs. They didn't show 16 is fine for 32 GPUs, and they didn't claim a fixed small number holds if you scale to 128 or 256 GPUs. Read it as 'don't over-provision beyond your topology,' not 'translation caching barely matters.' 27 00:10:09,424 --> 00:10:23,449 [Hal Turing] Fair, but doesn't that undercut their own framing a bit? They title the finding around minimal temporal locality making bigger TLBs pointless — that reads like a general claim to me, not a per-topology one. 28 00:10:23,449 --> 00:11:16,774 [Dr. Ada Shannon] It's a fair tension in their wording, I'll give you that. But mechanically, their own data ties the sufficient L2 size directly to GPU count — that's the mechanism, not a coincidence. So I'd land on: the qualitative finding generalizes, don't over-provision, but '32 is enough' is tied to this specific 32-GPU setup, not a universal constant. Either way, we agree blindly oversizing the L2 is the wrong instinct. Which is also why the mitigations they propose target the miss itself, not TLB capacity. Two directions, neither implemented yet: fused pre-translation kernels, which bake a pre-translation request into the compute kernel right before the collective fires, so the page walk finishes while the GPU is still busy computing. And software-guided TLB prefetching — using static knowledge of buffer layouts, or runtime profiling of repeated patterns, to warm likely entries before the request even lands. 29 00:11:16,774 --> 00:11:39,549 [Hal Turing] Picking up on "tied to this specific baseline" — that's the thread I want to pull. That whole hierarchy is borrowed, remember. So before I trust the 1.4x number or "size the L2 to GPU count," I want to know: would a real vendor Link MMU with different sizing or walker parallelism actually change the conclusion, or just relabel the same curve? 30 00:11:39,549 --> 00:12:02,724 [Dr. Ada Shannon] Right, and they say as much themselves. But the mechanism isn't "the L1 has 32 entries," it's "a cold destination pays a full page walk, and the L2 only needs to cover as many pages as there are GPUs hitting it." Change the table and you shift the exact 1.4x and the exact "32 is enough" threshold — those are baseline-dependent and illustrative. The cold-miss dominance and working-set-bound sizing, though, are architectural, not artifacts of this particular config. 31 00:12:02,724 --> 00:12:25,599 [Hal Turing] I'll grant the mechanism. But every experiment runs one workload — MSCCLang's all-pairs AllToAll — explicitly modeled assuming zero cache hits, zero temporal locality. Real NCCL or RCCL doesn't look like that, and neither does the same collective firing over and over across training steps with the same routing buffers recurring. 32 00:12:25,599 --> 00:12:54,425 [Dr. Ada Shannon] Oh — hold on, let me jump on that, because it cuts both ways. Repeated calls would claw back some cold-miss penalty since warm entries persist. But look at Tutel — Hwang, Cui, Xiong, Yang, Liu, Hu, Wang and colleagues out of Microsoft, 2023 — MoE dispatch routes a different, irregular token count per expert on every call. With a fixed 2MB page and shifting buffer boundaries, a destination could juggle more than one live page per GPU. That's the exact assumption "modest L2 suffices" rests on, and it's never tested here. 33 00:12:54,425 --> 00:13:14,900 [Hal Turing] And scale compounds that. They test 8 to 64 GPUs; UALink pods go to 1,024, and their own logic says working set scales with GPU count. I think that's more than a footnote, Ada — shouldn't we know whether "size the L2 to the pod" holds anywhere near the fabric's real design point? 34 00:13:14,900 --> 00:13:31,126 [Dr. Ada Shannon] I don't think it's that dramatic, honestly. The relationship they show is linear — one live page per GPU, full stop. Scale to 1,024 and the prescription is just "use a bigger L2" — same rule, bigger number, not a new phenomenon. 35 00:13:31,126 --> 00:13:56,526 [Hal Turing] I actually disagree with you there, Ada. L2 capacity might scale linearly, but the page walker doesn't — a hundred parallel PTWs are shared across all UALink traffic at the destination. At 1,024 GPUs all cold-missing during warm-up, that shared walker pool is the contention point, not TLB size. Linear TLB math says nothing about whether the walker queue blows up. 36 00:13:56,526 --> 00:14:32,876 [Dr. Ada Shannon] ...Fair, that's a real distinction — TLB sizing and walker contention are separate bottlenecks, and they only stress-tested the first at scale. I'll concede the hundred-walker pool at 1,024 GPUs is untested, not disproven. Also worth noting: they only simulate UALink — NVLink gets equal billing in the intro as sharing this problem, but it's never modeled; NVSwitch's topology and whatever undisclosed translation hardware NVIDIA runs could behave differently. And absolute latencies here inherit whatever validation ASTRA-sim itself carries — directional, not gospel. 37 00:14:32,876 --> 00:15:05,176 [Hal Turing] There's a gap they never touch, too: access control. NPA-to-SPA translation is also the permission boundary between OS domains in a shared pod, and none of this models isolation checks — a TLB right-sized for streaming throughput might be undersized once you add tenant metadata. And every experiment here is collective traffic. Point-to-point transfers — KV-cache migration between prefill and decode nodes — aren't GPU-count-bounded and don't stream the same way. Their central recommendation may just not travel there. 38 00:15:05,176 --> 00:15:55,851 [Dr. Ada Shannon] Which is really the scope problem in one sentence. They call this "the first systematic study," spanning NVLink and UALink with two concrete fixes — but what's delivered is one fabric, one synthetic workload, 64 GPUs max, on a borrowed baseline, and the fixes are never implemented, only proposed. There's precedent they'd work — Punniyamurthy, Hamidouche and Beckmann out of AMD Research, the same senior author here, published fused compute-collective kernels in 2024 — so it's plausible, just not evaluated. And the 20%-of-runtime inference stat motivating the urgency comes from TokenWeave — Gond, Kwatra and Ramjee, Microsoft Research India, 2025 — worth checking it actually lines up with the 1MB-scale collectives studied here, not inference generally. 39 00:15:55,851 --> 00:16:33,151 [Hal Turing] So here's where I land: cold misses crushing small collectives, and TLB capacity tracking GPU count instead of ballooning — that's a genuinely useful first-order result. But treat the 1.4x, the 32-entry threshold, and "don't over-provision" as illustrative, not gospel, past 64 GPUs, under real MoE traffic, or for point-to-point KV migration. Keep Link TLBs modest, but put the real effort into cold-miss mitigation — nobody's built that yet. That's it for this one. Thanks for listening, everybody — see you next time.