Serving MoE Models with Disaggregated Expert Parallelism

MegaScale-Infer treats attention and expert feed-forward execution as two different serving problems. This page visualizes the paper’s core claim: sparse MoE math can still waste GPU time unless routing, communication, and decode scheduling are redesigned together.

Visual focus: decode bottlenecks
Interactive modes: baseline vs DEP
Companion link: published viz
Paper MegaScale-Infer (2025)
Focus MoE decode serving
Key idea attention / experts split
Debate throughput vs TTFT

Split The Serving Path

Attention GPUs stay busy hauling KV-cache-heavy decode work while expert GPUs specialize in sparse FFN execution. Hover over stages to inspect the pressure points that make single-stack MoE serving under-utilized.

attention path router + expert path communication / pipeline handoff
The visual emphasis is not model semantics but resource asymmetry: decode attention is dominated by KV movement, while experts are dominated by uneven token arrivals and memory-bound FFNs.

Routing Skew Heatmap

MoE efficiency depends on how evenly tokens reach experts. Toggle between prefill and decode to see why smaller decode microbatches amplify underfilled experts and uneven loads.

Mock data is shaped to illustrate the paper’s story: larger prefill batches smooth routing, while decode batches create sharper hotspots, cold experts, and more idle memory bandwidth.

Ping-Pong Decode Timeline

Step through the microbatch pipeline. The baseline runs attention then experts in serial islands; the disaggregated schedule overlaps them so both GPU pools stay occupied more often.

This adapts pipeline intuition from GPipe and PipeDream, but here the split is by subsystem behavior rather than by stacking model layers end to end.

Benchmarks And The Skeptical View

Switch between a decode-heavy reading of the results and an end-to-end skeptical lens. The mock bars keep the paper’s qualitative ranking but make the tradeoff visible: throughput wins do not automatically map to first-token wins.

Illustrative numbers mirror the transcript’s framing: strong decode throughput, good cost efficiency, and more ambiguity once TTFT and shorter outputs matter.

References

Compact pointers for the papers and podcast episodes explicitly grounding this visualization. Transcript arXiv extraction found no additional IDs beyond 2504.02263.

MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
Ruidong Zhu et al., 2025 · arXiv:2504.02263
GPipe and PipeDream
Pipeline scheduling background for overlapping microbatches across stages.
Infinite-LLM, DistServe, Splitwise
Prior phase-splitting and attention-centric serving designs that motivate subsystem disaggregation.
MoE-Lightning, Toward Efficient Inference for MoE, AdapMoE, HybriMoE, Oracle-MoE
Related MoE inference work on memory pressure, locality, gating, and hybrid scheduling.
AI Post Transformers: JANUS for Scalable MoE Inference
AI Post Transformers: Splitwise, Prefill-as-a-Service, NanoFlow, LAPS
Prior podcast context on phase splitting, KV handling, and workload-shaped serving optima.