1 00:00:01,000 --> 00:00:44,096 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper is Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism, by Sajal Dash and one co-author, Feiyi Wang -- so a two-person team -- out of Oak Ridge National Laboratory, posted to arXiv on May 6th, 2026. Ada, the number that jumps off the page here: X-MoE, basically the current leading framework for this kind of training, only manages about 5.23 percent GPU utilization on a 545-billion-parameter model. 2 00:00:44,096 --> 00:01:29,003 [Dr. Ada Shannon] Five percent. Sit with that before we get into how Piper claims to fix it. That means on a national-lab supercomputer, roughly nineteen out of every twenty GPU-hours you're paying for are burned on waiting -- network contention, idle GPUs -- not on actual math. That's really the whole question this paper is chasing: can a resource model that combines pipeline parallelism with expert parallelism close the gap between the FLOPs you're paying for and what you actually get out of a messy, Dragonfly-topology HPC network? Piper claims two to three-and-a-half times that utilization. If that holds up, it's the difference between trillion-parameter training being a paper exercise and something Oak Ridge can actually run on Frontier without setting money on fire. 3 00:01:29,003 --> 00:01:47,161 [Hal Turing] So why is Mixture-of-Experts training specifically so much worse on a shared HPC platform than it'd be on a purpose-built AI cluster? And while you're at it, let's get listeners grounded -- what actually is MoE, structurally, versus the transformers we usually talk about here? 4 00:01:47,161 --> 00:02:36,759 [Dr. Ada Shannon] Three compounding problems: memory footprint from storing way more parameters than are active per token, load imbalance across experts, and non-uniform networks where bandwidth depends heavily on distance. Now, structurally: a normal transformer runs every token through the same dense feed-forward block. MoE swaps that for a bank of many smaller expert FFNs, with a learned router sending each token to just one or two. GShard, from Lepikhin and colleagues at Google in 2020, and Switch Transformer from Fedus, Zoph, and Shazeer in 2021, used a handful of large experts. The newer wave, pioneered by DeepSeek-MoE, goes fine-grained -- DeepSeek-V3 alone routes among 256 experts -- trading tensor-parallel complexity for a pile of small, awkward matrix multiplies. 5 00:02:36,759 --> 00:02:44,561 [Hal Turing] Okay, so once those experts are scattered across GPUs, how does a token actually find the right one, and why does that turn into a bottleneck at scale? 6 00:02:44,561 --> 00:03:12,703 [Dr. Ada Shannon] That's expert parallelism, the standard approach since GShard: each GPU owns a subset of experts, and every token gets physically shipped to whichever GPU holds its assigned expert. That's an all-to-all communication step, and it fires four times per MoE layer -- twice forward, twice backward. On a uniform fabric that's manageable. On Frontier's network, it isn't, which is exactly why Piper reaches for a tool nobody's applied here before. 7 00:03:12,703 --> 00:03:23,803 [Hal Turing] Oh wait, hold on -- pipeline parallelism? Isn't that the thing for splitting layers across GPUs in dense models, completely separate from how experts get distributed? 8 00:03:23,803 --> 00:04:39,685 [Dr. Ada Shannon] Exactly the tool, repurposed. Pipeline parallelism normally slices a model's layers across devices and streams micro-batches through with the 1F1B schedule -- built for dense models, where communication happens between layers, not inside them. Piper's insight is using that same layer-splitting to also cage the expert-parallel all-to-all inside a small, local group of GPUs. That locality question is exactly why Frontier's shape matters: it uses a Dragonfly topology, from Kim, Dally, Scott, and Abts' 2008 ISCA paper -- fast links within a node or switch group, much sparser links between groups. And it's also why static expert placement causes trouble: as training runs, some experts get more tokens than others, and without a way to move experts to match the load, GPUs sit unevenly loaded for the whole run. That's a physical fix, by the way, not a routing tweak -- unlike the auxiliary load-balancing losses everyone already bakes into the router, this actually moves expert weights between GPUs rather than penalizing the math. That's the migration problem we'll dig into properly once we get into Piper's design. 9 00:04:39,685 --> 00:05:22,735 [Dr. Ada Shannon] Right -- Piper organizes the whole GPU allocation into a PP-by-EP grid. You get PP pipeline stages, and each stage is staffed by only EP GPUs, which handle one slice of layers via expert-data parallelism. The trick is that the expert-parallel all-to-all -- the expensive part -- only ever has to happen inside that small EP group, ideally GPUs sitting on the same node or at most a single-hop switch group on Frontier. Compare that to X-MoE or DeepSpeed-MoE, where experts get scattered across the entire job and every dispatch and combine call has to reach across the whole allocation. Piper just refuses to let that communication group grow past what the fast interconnect can handle. 10 00:05:22,735 --> 00:05:38,060 [Hal Turing] Okay, but who decides what PP and EP should actually be for a given model and node count? That sounds like a huge combinatorial search -- pipeline depth, expert count per GPU, microbatch size, whether to checkpoint activations. 11 00:05:38,060 --> 00:06:20,971 [Dr. Ada Shannon] That's exactly what the resource model is for. Before running a single training step, Piper plugs the model's hidden dimension, layer count, expert count, and batch size into analytical equations for memory and communication cost, then filters out any configuration that would OOM or push expert-parallel traffic outside the fast-interconnect domain. What survives that filter gets ranked by a micro-benchmarking pass -- they profile attention throughput and expert GEMM throughput separately on a single GPU, find the batch-size sweet spot for each, then feed measured all-to-all and point-to-point bandwidth into an MFU estimator. So it's math to prune the search space down to a handful of candidates, then real hardware measurements to pick the winner among those. 12 00:06:20,971 --> 00:06:29,841 [Hal Turing] So the all-to-all itself -- even confined to a small group -- still has to cross node boundaries most of the time, right? That's where HALO comes in? 13 00:06:29,841 --> 00:07:07,132 [Dr. Ada Shannon] Right. HALO is their Dragonfly-aware all-to-all, and it's structured as three phases with a specific dependency chain: intra-node extraction, inter-node exchange, then intra-node redistribution. Phase one is fully independent of the other two, so they launch it concurrently with phase two or three and hide its latency behind the inter-node transfer. It also assigns GPUs to NICs so all four NICs on a node saturate simultaneously instead of contending, and where the scheduler allows it, it keeps node allocation inside a single rack to dodge the slowest inter-rack Dragonfly links entirely. 14 00:07:07,132 --> 00:07:10,522 [Hal Turing] And does that actually beat just calling RCCL's flat all-to-all? 15 00:07:10,522 --> 00:07:46,467 [Dr. Ada Shannon] Substantially, once you're past a certain scale. Against torch's RCCL-backed all_to_all, HALO is roughly on par below sixteen nodes -- everything still fits inside one switch group, so flat all-to-all already saturates the link. But past sixteen nodes, where inter-rack traffic starts dominating under the flat approach, HALO pulls ahead one-point-one to nine times depending on message size and node count, with the biggest wins around thirty-two nodes. That crossover point is basically the paper's proof that topology-awareness only pays off once you're forced to cross the slow part of the fabric. 16 00:07:46,467 --> 00:07:59,470 [Hal Turing] Sorry, hold on -- back up a second, because I want to make sure I've got the migration piece straight before we move to results. Why does expert placement even need to move around mid-training? I thought experts just sit where they're assigned. 17 00:07:59,470 --> 00:09:05,600 [Dr. Ada Shannon] Because the router's preferences drift. Early on, all experts look roughly equal, but tiny random differences make some slightly more attractive, they get more tokens, more gradient updates, get better at their job, and attract even more tokens -- a feedback loop the paper calls expert collapse. Auxiliary load-balancing losses only smooth this at the routing level; they don't fix the fact that some GPU is now sitting there doing four times the work of its neighbor. Piper's answer is to physically swap experts between GPUs in the same layer using a hill-climbing algorithm: find the most-loaded and least-loaded GPU, test candidate swaps, take the one that shrinks the gap the most, repeat up to a hundred iterations. Because migration only moves experts within that same fast-interconnect neighborhood, Table IV's worst-case per-GPU migration latency comes out under five percent of total training time even for something like DeepSeek-V3's configuration. Worth flagging plainly -- the actual load-distribution measurement behind this, Figure 9, is on a fairly small model: 350 million parameters, 16 experts. 18 00:09:05,600 --> 00:09:11,684 [Hal Turing] Got it -- so how does all this actually translate into throughput once you put a real model through it? 19 00:09:11,684 --> 00:10:07,226 [Dr. Ada Shannon] Single-layer ceiling first: they train one layer of each SOTA model on a single Frontier node and get anywhere from about 78 TFLOPS for DeepSeek-V2 up to 129 for Mixtral 8x22B, averaging around 102. That's the upper bound before pipeline bubbles enter the picture. Full-model runs with activation checkpointing land Piper's MFU between 19.8 and 53.8 percent depending on the model -- Mixtral's coarse-grained experts hit the high end, fine-grained models like DeepSeek-V2 sit lower. And against the field -- Tutel from Hwang et al. at Microsoft, MLSys 2023; DeepSpeed-MoE, Rajbhandari et al., ICML 2022; DeepSpeed-TED, the Microsoft DeepSpeed team's 2022 report; and X-MoE -- Piper delivers two to three-point-six times the throughput while using a fraction of the GPU count those frameworks needed for the same model size. 20 00:10:07,226 --> 00:10:10,941 [Hal Turing] And the trillion-parameter number everyone's going to remember from the abstract? 21 00:10:10,941 --> 00:10:44,378 [Dr. Ada Shannon] They scale a 10-billion-parameter dense base out by expert count -- 16 experts up through 256 -- and train an 862-billion-parameter model on 512 GPUs at 39.38 TFLOPS, then a 1.7-trillion-parameter model on 1024 GPUs at 33 TFLOPS. Going from 64 GPUs to 1024, weak-scaling efficiency holds at 73 percent against the ideal throughput line. Those are the headline scale numbers -- worth sitting with them for a second before we get into how solid the evidence behind them actually is. 22 00:10:44,378 --> 00:11:08,805 [Hal Turing] Seventy-three percent weak-scaling efficiency is genuinely impressive, Ada, but I want to go back to something that's been nagging me since you mentioned expert migration. That whole mechanism is the fix for load imbalance, right? So was it actually running during those 862-billion and 1.7-trillion-parameter throughput runs, or is the sub-5%-overhead number just a paper calculation? 23 00:11:08,805 --> 00:11:39,827 [Dr. Ada Shannon] Good catch, because when you actually dig into Section VII-D, migration is never mentioned as enabled. The 5% figure traces back to Table IV, which is a worst-case per-GPU latency estimate for a full expert reshuffle at various model sizes, computed from bandwidth assumptions, not measured end-to-end during a trillion-parameter training step. So what gets marketed as 'migration costs under 5% of training time' is an extrapolated ceiling from a static latency table, not something they clocked while actually running the 1.7T configuration. 24 00:11:39,827 --> 00:11:48,837 [Hal Turing] And that's the only place migration gets validated with real data at all -- which puts a five-orders-of-magnitude gap between what's demonstrated and what's claimed. 25 00:11:48,837 --> 00:12:23,156 [Dr. Ada Shannon] Right, and it gets worse. Section VI states outright that 'even with the expert collapse, the final model performance does not suffer' -- no loss curve, no perplexity comparison against a load-balanced baseline, no downstream eval. That's a strong claim stated as fact. It also runs counter to what the field generally finds: expert collapse degrades specialization, which is exactly why DeepSeek-V3's bias-adjustment balancing and expert-choice routing exist in the first place. Making that claim without evidence is the paper's weakest moment. 26 00:12:23,156 --> 00:12:40,385 [Hal Turing] Wait, sorry -- hold on, because that connects directly to my next question. If the trillion-parameter runs in Section VII-D only report TFLOPS and weak-scaling efficiency, are we even sure those are converging training runs at all, or just forward-backward throughput demos on whatever data happened to be in the loader? 27 00:12:40,385 --> 00:13:27,289 [Dr. Ada Shannon] Exactly the right question, and the paper gives you no way to answer it. There's no loss value reported anywhere in that section. What Piper demonstrates convincingly is that it can execute training steps at a given hardware utilization on Frontier -- that's a real systems result. What it does not demonstrate is that those steps produce a model converging toward anything useful. The paper conflates 'training a trillion-parameter model' with 'achieving high MFU on a trillion-parameter model's forward and backward pass,' and those are different claims. On top of that, their own reference list flags X-MoE's arXiv ID as incorrect -- it points to an unrelated paper -- so even the 5.23% MFU baseline everyone's measuring against needs independent verification. 28 00:13:27,289 --> 00:13:41,593 [Hal Turing] That's a rough one to catch in your own bibliography. So practically -- for an HPC center sitting on AMD GPUs and a Dragonfly network instead of InfiniBand, what should they take from this paper? 29 00:13:41,593 --> 00:14:37,739 [Dr. Ada Shannon] The systems contribution is real and worth adopting: confining expert-parallel communication to a topologically local group via pipeline parallelism, and HALO's rack-aware three-phase all-to-all, are genuinely useful ideas for anyone stuck with non-uniform interconnects. Where I'd push back is the load-balancing story. FlexMoE, from Nie, Miao, Wang, Yang, Xue, Ma, Cao, and Cui at Peking University and Microsoft, published at SIGMOD 2023, already does dynamic device placement for exactly this problem, and Piper's related-work section never engages it. SmartMoE, from Zhai, He, Ma, Zong, Zhang, and Zhai at Tsinghua, USENIX ATC 2023, builds a nearly identical offline-plus-online parallelization search to Piper's resource model, and again, no comparison. Piper's throughput numbers are solid; its load-balancing and trillion-scale claims are asserted, not shown. 30 00:14:37,739 --> 00:15:01,330 [Hal Turing] So net assessment: strong communication and pipelining system, but the headline trillion-parameter and load-balancing claims are riding on evidence that doesn't match their scale. That's Piper -- worth stealing the HALO and pipeline-mesh ideas, worth being skeptical of everything past the throughput chart. Thanks for breaking this one down, Ada, and thanks everyone for listening -- we'll catch you next time. 31 00:15:01,330 --> 00:15:02,491 [Dr. Ada Shannon] Take care, everyone.