MegaScale-Infer treats attention and expert feed-forward execution as two different serving problems. This page visualizes the paper’s core claim: sparse MoE math can still waste GPU time unless routing, communication, and decode scheduling are redesigned together.
Attention GPUs stay busy hauling KV-cache-heavy decode work while expert GPUs specialize in sparse FFN execution. Hover over stages to inspect the pressure points that make single-stack MoE serving under-utilized.
MoE efficiency depends on how evenly tokens reach experts. Toggle between prefill and decode to see why smaller decode microbatches amplify underfilled experts and uneven loads.
Step through the microbatch pipeline. The baseline runs attention then experts in serial islands; the disaggregated schedule overlaps them so both GPU pools stay occupied more often.
Switch between a decode-heavy reading of the results and an end-to-end skeptical lens. The mock bars keep the paper’s qualitative ranking but make the tradeoff visible: throughput wins do not automatically map to first-token wins.
Compact pointers for the papers and podcast episodes explicitly grounding this visualization. Transcript arXiv extraction found no additional IDs beyond 2504.02263.