1 00:00:01,000 --> 00:00:36,899 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're covering a paper called Coordinated Scheduling for MoE LLM Serving — the system they build is called Gimbal. That's Yifan Sun et al., eight authors total, out of the University of Melbourne, Monash University, and the Shenzhen Institutes of Advanced Technology at the Chinese Academy of Sciences. It landed on arXiv June 13th, 2026. Ada, this is squarely your territory — serving infrastructure for sparse models. 2 00:00:36,899 --> 00:00:56,899 [Dr. Ada Shannon] What sold me on this one is they didn't just tune heuristics until the benchmark numbers looked good — they formally frame DP-engine dispatch and expert placement as one coupled optimization problem before they design the fix. That kind of grounding usually travels better than a pile of tuned knobs. Whether it actually holds up outside their test setup is a fair question — we'll get there. 3 00:00:56,899 --> 00:01:18,849 [Hal Turing] For anyone new to serving systems: a frontend router hands incoming requests off to one of several backend engine replicas. Those dispatch policies — round robin, pick-the-shortest-queue — were built for dense models, where every token costs roughly the same. Why does that assumption break the moment you bring Mixture-of-Experts into the picture? 4 00:01:18,849 --> 00:01:58,949 [Dr. Ada Shannon] Because requests stop being interchangeable, on two fronts. A Mixture-of-Experts model doesn't fire every parameter for every token — each layer has a bank of expert sub-networks, and a lightweight router activates only a handful, the top-k, per token. That's how you get huge model capacity without paying huge compute per token — DeepSeek's models are the obvious example. On the serving side, Data Parallelism, DP, just replicates the whole engine — separate request queues, separate KV caches. Expert Parallelism, EP, scatters the experts themselves across GPUs, so tokens have to physically travel — all-to-all communication — to wherever their chosen expert lives. 5 00:01:58,949 --> 00:02:16,750 [Hal Turing] So the router is making a decision the frontend scheduler can't see coming — it picks which engine a request enters through with zero idea that, three layers later, that request triggers a scramble across GPUs it doesn't own. Is expert locality the fix, then? 6 00:02:16,750 --> 00:02:45,500 [Dr. Ada Shannon] Locality is part of it — place an expert near whichever DP engine's tokens hit it most often, so the traffic stays local instead of paying that cross-group communication tax. But that only works if placement actually knows the source of the traffic, and that's where coordinated cross-level scheduling comes in: feed backend pressure — KV-cache usage, queue depth, expert load — into the frontend's dispatch decision, and feed observed request-to-expert traffic back into placement. That's workload-aware scheduling in a nutshell — react to live state, not a static counter. 7 00:02:45,500 --> 00:03:02,574 [Hal Turing] They've got a clean example for this. Picture two DP engines, each with exactly one request in flight — so by request count, they look perfectly balanced. One request is 200 tokens, the other is 2,000. On paper, that's a wash. In practice— 8 00:03:02,574 --> 00:03:25,074 [Dr. Ada Shannon] Oh, I know this one — not even close. Same Qwen3 setup on an A100: the 200-token request costs 18.75 megabytes of KV cache and hits about 86 milliseconds time-to-first-token. The 2K-token request costs 187.5 megabytes — ten times the memory — and TTFT climbs to nearly 174 milliseconds. Identical request count, wildly different pressure. 9 00:03:25,074 --> 00:03:41,349 [Hal Turing] And that's just compute and memory pressure — one axis of imbalance. Inside the model, routing itself isn't remotely uniform either, with some experts getting hammered while others barely fire at all. What did the profiling actually turn up? 10 00:03:41,349 --> 00:04:10,474 [Dr. Ada Shannon] They served a thousand requests on Qwen3-80B and traced activation per layer — the hottest experts absorb way more than the layer average, classic hotspot behavior. But the sharper finding is source-dependent: under the current placement, 83.4% of one layer's traffic originating from DP engine zero routes to experts sitting in a remote DP group. Not local — remote. Most of that traffic is paying full cross-GPU communication cost that better placement could avoid entirely. 11 00:04:10,474 --> 00:04:25,225 [Hal Turing] Okay, but honestly — isn't the fix here just 'move the expert closer to where its traffic comes from'? That sounds less like a research contribution and more like something a decent systems engineer sketches on a whiteboard. 12 00:04:25,225 --> 00:04:46,550 [Dr. Ada Shannon] I actually disagree with you there, Hal. Deciding locality matters is the easy part — the hard part is you can't do source-aware placement without solving migration cost, and you can't fix dispatch without backend pressure that shifts every few hundred milliseconds. That's exactly why everyone before this either ignored source information entirely or rebalanced experts offline on a fixed schedule. 13 00:04:46,550 --> 00:05:02,025 [Hal Turing] Sure, but doesn't offline rebalancing already capture most of that value for a fraction of the engineering cost? At some point you're chasing diminishing returns on a problem EPLB-style approaches already mostly solve. 14 00:05:02,025 --> 00:05:41,650 [Dr. Ada Shannon] Maybe on the expert side alone — but offline rebalancing is blind to the DP-scheduling half of this entirely. It looks at expert placement maybe once every few minutes, with zero visibility into a request landing on an engine sitting at ninety percent KV usage while its neighbor idles at twenty. Those two failure modes compound each other — fixing them separately means solving half the equation twice while missing the interaction. The numbers settle it for me: 42.9% lower time-to-first-token, 33.3% lower time-per-output-token, against vLLM. That's not diminishing returns — the coordination between the two halves is doing the work, and that's exactly what we should dig into next. 15 00:05:41,650 --> 00:05:55,250 [Hal Turing] Alright, alright, you've got the receipts. So walk me through the mechanics — 'coordinate the scheduler and the placement engine' is easy to say on a whiteboard, harder to build without the two fighting for control. 16 00:05:55,250 --> 00:06:18,600 [Dr. Ada Shannon] Two components sharing a nervous system. A fine-grained DP-engine scheduler sits in front of the request pool, a source-aware expert load balancer sits inside the MoE execution path, and two feedback signals cross between them: backend MoE pressure flows up into the DP scheduler's dispatch decision, and observed source-DP-to-expert traffic flows into where experts get physically placed. Neither half is blind to the other anymore. 17 00:06:18,600 --> 00:06:30,150 [Hal Turing] Before a request even gets assigned, then, the scheduler needs a read on engine state. What's actually in that signal — because 'engine pressure' could mean a dozen things. 18 00:06:30,150 --> 00:06:56,300 [Dr. Ada Shannon] Four numbers per engine, reported asynchronously so collection never blocks the request path: remaining prefill tokens still owed by running requests, waiting prefill tokens in the local queue, KV-cache usage, and backend MoE expert pressure. Splitting remaining-prefill from waiting-prefill matters because of chunked prefill — prompt length at admission time isn't what's actually left to compute once continuous batching spreads that work across iterations. Each engine phones this home on its own schedule. 19 00:06:56,300 --> 00:07:06,075 [Hal Turing] And if you're dispatching faster than traces update, doesn't the scheduler just keep hammering whichever engine looked best on the last stale snapshot? 20 00:07:06,075 --> 00:07:35,300 [Dr. Ada Shannon] Two safeguards against exactly that. KV cache gets priority — if usage is high with a clear gap to the least-loaded engine, Gimbal routes there directly, no scoring, because exhausting KV cache triggers preemption and recomputation, a far worse outcome than an uneven score. Otherwise it's a composite score plus a compensation term: after every dispatch, Gimbal adds a temporary pressure estimate to whichever engine just got the request, so it doesn't pile on before the next refresh. And when two scores are close enough to be noise, it just falls back to ordered dispatch. 21 00:07:35,300 --> 00:07:51,550 [Hal Turing] Oh wait, hold on — and once a request lands in an engine's local queue, that's the shortest-job-first-with-aging thing, right? Order by prefill length so short requests don't queue behind a monster prompt, with an escape hatch so nothing starves. 22 00:07:51,550 --> 00:08:10,125 [Dr. Ada Shannon] Exactly, deliberately cheap — no output-length prediction, just prefill token count as the proxy, with an aging threshold promoting stale requests to the front. The more interesting half is the expert side: a layer-wise source-DP-to-expert matrix, tracking not just which experts are hot, but who's making them hot. 23 00:08:10,125 --> 00:08:21,400 [Hal Turing] That distinction feels like the whole ballgame — EPLB-style balancing only sees the aggregate count. What does placement actually do with that extra dimension? 24 00:08:21,400 --> 00:08:59,775 [Dr. Ada Shannon] It builds a placement objective around three goals at once: balance expert load across ranks, minimize source-aware communication using that traffic matrix, and limit migration cost, since moving experts mid-serving costs roughly a second per full rearrangement on their testbed. Solve that exactly and you want the full mixed-integer nonlinear program — except solving it for a 48-layer Qwen3-30B model takes about fifteen seconds, an eternity on a serving critical path. So the MINLP runs offline as a calibration target, and a much cheaper greedy heuristic runs online, tuned to track what the exact solver would've picked. 25 00:08:59,775 --> 00:09:09,850 [Hal Turing] So the greedy version is 'the MINLP said roughly this, now go do it fast.' Fair — but does the shortcut hold up? Give me the headline numbers. 26 00:09:09,850 --> 00:09:47,676 [Dr. Ada Shannon] 42.9% lower average TTFT and 33.3% lower average TPOT against vLLM, across every request rate and workload distribution tested. Throughput went up too, 3% at high load, so this isn't latency bought at throughput's expense. P99 TTFT dropped 44.3%, which arguably matters more than the average since tail latency is what breaks SLAs. And the gains aren't flat — TTFT improvement climbs from around 33% at the lowest request rate up past 48% at the highest. This is a system that gets more valuable exactly when you're under the most pressure. 27 00:09:47,676 --> 00:10:01,476 [Hal Turing] That growth-with-load curve is what actually convinces me. But did they prove the coordination itself is doing the work, or is this just two good ideas stacked together and dressed up as something more unified? 28 00:10:01,476 --> 00:10:56,151 [Dr. Ada Shannon] That's exactly what the ablation is for, and it's the paper's central empirical claim. DP-only scheduling alone: 25% TTFT improvement. EP-only placement alone: 26%. Stack both independently, zero collaboration, and you get to roughly 30% — respectable, basically additive. Turn on actual coordination — feeding backend MoE pressure back into the DP scheduler instead of running placement and dispatch as parallel tracks — and TTFT improvement jumps to 41%, TPOT to 32%: another 16.5% off TTFT and 6.5% off TPOT on top of the no-collaboration stack, holding across every request rate and distribution tested. That gap between stacked and coordinated is the number that justifies the architecture — the interaction itself is buying real latency, not just two decent ideas sharing a codebase. 29 00:10:56,151 --> 00:11:27,426 [Hal Turing] I actually disagree with you there, Ada. Sixteen and six percent is real, but look at the cost. Gimbal-DP alone gets 25.1% TTFT, Gimbal-EP alone gets 26.2%, and just stacking both uncoordinated gets you nearly 30. Wiring backend expert pressure directly into frontend dispatch is a much tighter coupling — more state to keep consistent, more ways for it to go wrong. For a few extra points, is that operational complexity actually worth it? 30 00:11:27,426 --> 00:11:49,201 [Dr. Ada Shannon] No no no — wrong comparison. P99 TTFT is down 44.3% against vLLM, and coordination is exactly where tail latency lives. The requests that spike are the ones landing on an engine whose co-located expert ranks just got hammered. Complexity that specifically kills your tail is complexity worth paying for. Users don't feel the median, they feel the one four-second request. 31 00:11:49,201 --> 00:12:27,751 [Hal Turing] Fair — tail latency as the justification lands better than just citing the average. But that coupling worries me once you leave this testbed. The whole evaluation is one node, four H100s, DP equals two. The intro motivates itself with DeepSeek-V3-style deployments — dozens of DP replicas, multi-node EP across hundreds of GPUs. The async trace needs a compensation term because reporting is asynchronous. At two engines over NVLink, fine. At sixteen or thirty-two engines with real network round-trips, does that same fixed term still keep the scheduler from oscillating? 32 00:12:27,751 --> 00:12:50,751 [Dr. Ada Shannon] Honest gap — the conclusion says outright, 'evaluated on a single-node deployment.' But mechanically I'd bet it gets worse, not better: that compensation term is a fixed estimate calibrated against sub-millisecond round-trips. Multi-node collection means longer, more variable delays, so the gap between what the scheduler believes and what's true widens — nothing adapts the term to round-trip time. More oscillation risk exactly where they claim coordination matters most. 33 00:12:50,751 --> 00:13:22,701 [Hal Turing] Oh wait, hold on — it's not just node count, it's the model too. They benchmark Qwen3-30B-A3B, about three billion active params, but the hotspot and traffic-skew figures that motivate the whole paper run on Qwen3-80B-INT4 — a different, bigger model. And the intro keeps invoking DeepSeek-V2, V3, V4 as the real stakes. Nobody measures the 42.9 and 33.3 percent numbers past thirty billion parameters. 34 00:13:22,701 --> 00:14:03,776 [Dr. Ada Shannon] Worth reasoning through which way that cuts, not just flagging it. All-to-all communication overhead scales with EP degree — more experts, more GPUs, more cross-rank hops for a remote-expert token. Hotspot severity likely worsens too as activation spreads over more experts per layer. So I'd bet gains grow at scale, not shrink — but that's a bet, not a measurement. And it rhymes with Section 5.3: the full MINLP takes fifteen seconds for this 48-layer model, so they calibrate the greedy heuristic's weights against it once, offline, on this exact model and testbed — the same offline-profiling trap the paper spends a paragraph criticizing Sem-MoE, Li, Zhang, Wang, Chen, and Zheng, ICLR 2026, for. 35 00:14:03,776 --> 00:14:26,326 [Hal Turing] Right — isn't that exactly the trap? Figure 6 only shows the calibration holding on this one model and topology, over 80% placement match, comm cost within 0.6% of the MINLP reference. Change the expert count or GPU layout and there's no evidence those alpha-beta-gamma weights still track the offline optimum. 36 00:14:26,326 --> 00:14:47,801 [Dr. Ada Shannon] There's a real distinction, even if smaller than they'd like. MoETuner — that's Go and Mahajan, Georgia Tech, 2025 — computes a static placement once offline and never updates it. Gimbal's source-DP-to-expert matrix keeps updating online every window; only three scalar weights are fixed offline, not the placement itself. Inputs stay live even though the weighting doesn't — real difference, just narrower than the framing suggests. 37 00:14:47,801 --> 00:15:24,427 [Hal Turing] And there's a production wrinkle they sidestep — vLLM's prefix-cache matching gets disabled 'to reduce experimental bias.' But prompt-cache locality is exactly what Preble: Efficient Distributed Prompt Scheduling for LLM Serving, by Srivatsa, He, Abhyankar, Li, and Zhang out of UC San Diego, ICLR 2025, optimizes for — routing to wherever the matching KV prefix already lives. That's a different objective than lowest-pressure dispatch. Turn caching back on and the two signals compete for the same decision, and Gimbal never says how that gets resolved. 38 00:15:24,427 --> 00:16:00,852 [Dr. Ada Shannon] Interesting thread — Zhexiang Zhang and Adel Toosi, two co-authors on this very paper, also wrote JANUS: Disaggregating Attention and Experts for Scalable MoE Inference, 2025, disaggregating attention and experts onto separate GPU pools entirely — the opposite architectural bet from Gimbal's co-located design. The paper names that as a natural extension itself. So it depends on scale: two to four GPUs on one node, Gimbal's coordination is a real, measured win. Dozens of DP replicas chasing DeepSeek-V3 throughput, and the motivation has outrun the evidence. 39 00:16:00,852 --> 00:16:28,602 [Hal Turing] To be fair, they own that — the conclusion is upfront about single-node scope, naming more MoE models, multi-node deployments, and more adaptive placement as next steps. Bottom line: solid coordination engineering with real numbers on hardware most teams can rent — 42.9% TTFT, 33.3% TPOT. Just don't extrapolate onto a thirty-two-GPU DeepSeek-scale cluster until someone runs that experiment. 40 00:16:28,602 --> 00:16:37,077 [Dr. Ada Shannon] Agreed. Honestly scoped limitations and a clear list of what needs testing next — that's more than most systems papers give you. 41 00:16:37,077 --> 00:16:43,002 [Hal Turing] That's Gimbal. Thanks for listening to AI Post Transformers — catch you next time.