AI Post Transformers · Episode Companion

Coordinated MoE Scheduling: How Gimbal Fixes Expert Locality

arXiv 2606.15177 Yifan Sun, Zhexiang Zhang, et al. · 2026 Univ. of Melbourne · Monash · SIAT, CAS Testbed: 4×H100, DP=2 vs vLLM: −42.9% TTFT · −33.3% TPOT

Two decisions, coupled into one loop

Dense-model serving treats "which engine gets this request" as independent from "where do model weights live." MoE breaks that independence: a request's engine choice determines which experts it will collide with three layers later. Toggle below to see what changes when the two decisions are wired together.

4
async pressure signals feed the scheduler
2
feedback loops: pressure↑ and traffic→placement
1
coupled optimization problem, not two

Why request count lies to the scheduler

A Mixture-of-Experts layer activates only a top-k slice of experts per token (Data Parallel engines replicate the whole model; Expert Parallel groups scatter experts across GPUs, forcing all-to-all traffic whenever a token's chosen expert lives elsewhere). Round-robin and shortest-queue dispatch were built for dense models where every token costs the same — that assumption is exactly what MoE breaks. See the next tab for the numbers.

Same request count, wildly different pressure

Two DP engines, one request each — balanced by count. One prompt is 200 tokens, the other 2,000. Same Qwen3 setup, A100 GPU. Toggle the view to see what the scheduler sees vs. what's actually happening.

10×
KV-cache memory gap (18.75MB → 187.5MB)
2×
time-to-first-token gap (86ms → ~174ms)

Expert traffic is source-dependent, not just hot/cold

Profiling 1,000 requests on Qwen3-80B: rows are the DP engine a token originated from, columns are the 16 experts in one layer, grouped by which GPU group is "local" to which engine (thick separators). Hover a cell for its share of that engine's traffic.

83.4%
of DP-engine-0's traffic routes to remote experts
1000
requests profiled on Qwen3-80B-INT4

Dispatch: from signal to queue slot

Click through the five stages a request passes through before it lands in a local queue.

Placement: one objective, three pulls

Balancing load, minimizing source-aware communication, and limiting migration cost pull in different directions. Solving exactly needs a mixed-integer nonlinear program — too slow for the serving path, so it becomes an offline calibration target for a fast online heuristic. Hover any pipeline node for detail.

~15s
exact MINLP solve, 48-layer Qwen3-30B — too slow live
~1s
cost per full expert rearrangement
>80%
greedy placement match vs. MINLP reference (Fig. 6)

Headline numbers vs. vLLM

All bars are relative to vLLM = 100. Lower is better. Hover a bar for the exact figure.

+3%
throughput at high load — latency gains aren't bought from throughput
33%→48%
TTFT improvement grows as request rate climbs

Is the coordination itself doing the work?

DP-only and EP-only fixes stack to roughly additive gains. Turning on the feedback loop between them — not just running both — is what produces the jump.

Coordinated TPOT improvement: 32%, vs. TTFT-only figures charted above. P99 TTFT specifically drops 44.3% — tail latency is where the coordination pays off most, since spikes come from requests landing where co-located expert ranks just got hammered.

References

  1. Yifan Sun, Zhexiang Zhang, Jiantong Jiang, Gholamreza Haffari, Minxian Xu, Feng Liu, Rajkumar Buyya, Adel N. Toosi · 2026
  2. Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, Byung-Gon Chun · 2022
  3. Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, et al. · 2023
  4. DeepSeek-AI (Damai Dai, Wenfeng Liang, et al.) · 2024
  5. Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, et al. (Moonshot AI / Kimi) · 2024
  6. Vikranth Srivatsa, Zijian He, Reyna Abhyankar, Dongming Li, Yiying Zhang · ICLR 2025
  7. Seokjin Go, Divya Mahajan · 2025
  8. Yan Li, Zhenyu Zhang, Zhengang Wang, Pengfei Chen, Pengfei Zheng · ICLR 2026
  9. Jinghan Yao, Quentin Anthony, Aamir Shafi, Hari Subramoni, Dhabaleswar K. Panda · IPDPS 2024
  10. Zhexiang Zhang, Ye Wang, Yumiao Zhao, et al. · 2025