MegaScale-Infer: Disaggregating Experts for Faster MoE Serving

arXiv:2504.02263 ByteDance Seed · Peking University 20 authors · Apr 2025

A roofline-model teardown of why MoE's sparse routing wrecks GPU utilization during decode, and MegaScale-Infer's fix: physically splitting attention and expert computation onto independently scaled GPU pools so attention replicas can pool enough requests to keep experts saturated.

References