Dense-model serving treats "which engine gets this request" as independent from "where do model weights live." MoE breaks that independence: a request's engine choice determines which experts it will collide with three layers later. Toggle below to see what changes when the two decisions are wired together.
A Mixture-of-Experts layer activates only a top-k slice of experts per token (Data Parallel engines replicate the whole model; Expert Parallel groups scatter experts across GPUs, forcing all-to-all traffic whenever a token's chosen expert lives elsewhere). Round-robin and shortest-queue dispatch were built for dense models where every token costs the same — that assumption is exactly what MoE breaks. See the next tab for the numbers.
Two DP engines, one request each — balanced by count. One prompt is 200 tokens, the other 2,000. Same Qwen3 setup, A100 GPU. Toggle the view to see what the scheduler sees vs. what's actually happening.
Profiling 1,000 requests on Qwen3-80B: rows are the DP engine a token originated from, columns are the 16 experts in one layer, grouped by which GPU group is "local" to which engine (thick separators). Hover a cell for its share of that engine's traffic.
Click through the five stages a request passes through before it lands in a local queue.
Balancing load, minimizing source-aware communication, and limiting migration cost pull in different directions. Solving exactly needs a mixed-integer nonlinear program — too slow for the serving path, so it becomes an offline calibration target for a fast online heuristic. Hover any pipeline node for detail.
All bars are relative to vLLM = 100. Lower is better. Hover a bar for the exact figure.
DP-only and EP-only fixes stack to roughly additive gains. Turning on the feedback loop between them — not just running both — is what produces the jump.
Coordinated TPOT improvement: 32%, vs. TTFT-only figures charted above. P99 TTFT specifically drops 44.3% — tail latency is where the coordination pays off most, since spikes come from requests landing where co-located expert ranks just got hammered.