1 00:00:01,000 --> 00:00:53,149 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. And today we're digging into a paper out of Huawei called ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill. That's Weiwei Chen et al., eight authors total — Weiwei Chen, Shuang Chen, Lele Li, Qiang Hu, Han Li, Xin Ye, Ming Yan, and Zhibin Yu, all at Huawei. It landed on arXiv on June 21st, 2026. The question driving the whole thing is simple to state: can you physically split a mixture-of-experts model's attention stage and its expert stage onto separate hardware, connect them with communication that genuinely never blocks, and by doing that kill off the synchronization stalls that quietly wreck time-to-first-token in real production serving? 2 00:00:53,149 --> 00:01:21,099 [Dr. Ada Shannon] Here's what pulled me in: this isn't a paper that starts with a clever mechanism and backfills a justification. They characterize the problem first — real production traces, measuring exactly where the time goes, showing mathematically why the imbalance can't just be scheduled away. Only after that do they propose a fix. Most systems papers reverse that order, and you end up trusting the benchmark numbers more than the diagnosis. Here the diagnosis is the strongest part, and that's why it's worth our full attention today. 3 00:01:21,099 --> 00:01:42,974 [Hal Turing] Okay, before we get into what they built, let's set the stage — because I don't think we can assume the audience already knows why serving these models needs this hybrid parallelism setup in the first place. Ada, why can't you just parallelize a mixture-of-experts model the same way we'd parallelize a regular dense transformer and call it a day? 4 00:01:42,974 --> 00:02:42,574 [Dr. Ada Shannon] Because the whole point of MoE is that most parameters sit idle for any given token. Instead of one dense feed-forward block per layer, you've got a large pool of specialized expert sub-networks, and a router sends each token to just a handful — DeepSeek-V3 is the poster child, 671 billion total parameters but only 37 billion active per token. Since almost all of those parameters live in the experts, you spread them across accelerators with Expert Parallelism, EP, so no single device holds the whole pool. That creates a problem for attention, though: Tensor Parallelism, the usual way to speed up attention, gets communication-expensive past about 8-way, so it's capped. To keep attention matched to a wide EP degree, you're forced into Data Parallelism on top — DP equals EP divided by TP. The DeepSeek-V3 production instance runs TP=8, DP=4, EP=32 across 32 GPUs on 4 nodes, per the DeepSeek-V3 technical report, DeepSeek-AI, late 2024. 5 00:02:42,574 --> 00:03:00,000 [Hal Turing] Right, just so we're speaking the same language before we go further — can you break down prefill versus decode, and what TTFT actually measures? I've seen the term thrown around a lot, and I want us to nail down exactly what it means before we talk about stalls and latency. 6 00:03:00,000 --> 00:03:41,100 [Dr. Ada Shannon] Sure. Every request goes through two phases. Prefill processes the entire prompt at once and produces the first output token — it's compute-bound, because you're crunching every input token in parallel. Decode is what happens after: generating each subsequent token one at a time, autoregressively, and it's memory-bound, since you're mostly streaming weights and KV cache off memory for a tiny amount of compute per step. Time-to-First-Token, TTFT, is exactly what it sounds like — the latency from when a request lands to when that first token comes back. Because that's the number the user feels staring at a blank screen, its service target is usually just a few seconds. Decode's equivalent metric, time-per-output-token, gets tens of milliseconds instead, since it only has to keep pace with reading speed. 7 00:03:41,100 --> 00:03:59,450 [Hal Turing] Okay — so attention runs on Data Parallelism and the experts run on Expert Parallelism, both feeding into the same shared pool underneath. That sounds clean on paper. Where does it break down once you throw real, messy production traffic at it instead of a clean benchmark? 8 00:03:59,450 --> 00:04:20,750 [Dr. Ada Shannon] Because that shared expert pool forces a hard synchronization point. Every attention DP group processes its own batch independently, sure, but before and after every single MoE layer, all of those groups have to stop and rendezvous at a global barrier so the MoE stage can process everyone's tokens together. And a barrier is only ever as fast as its slowest participant— 9 00:04:20,750 --> 00:04:39,750 [Hal Turing] Oh, wait, wait — hold on, that's the MapReduce straggler problem, isn't it? Dean and Ghemawat, Google, 2004 — one slow reducer holds the whole job hostage even if every other node finished early. Is that basically the same failure mode, just relocated into model serving? 10 00:04:39,750 --> 00:05:23,525 [Dr. Ada Shannon] Exactly that, structurally. The paper calls it a straggler effect — the whole system stalls until the slowest attention DP group, or the most congested expert, finishes. What makes it nasty rather than just annoying: it's unavoidable by design, not a tuning problem. Arrivals are stochastic, sequence lengths are heavy-tailed — the Huawei Cloud trace ranges from 31 tokens to nearly 100,000, averaging around 5,000 — and attention cost scales quadratically with each sequence length, not the batch's total token count. So even perfectly balancing total tokens across DP groups still leaves compute time differing several-fold, since it depends on the sum of squares of each request's length, not the sum itself. No scheduling trick fixes that — it's baked into the math. 11 00:05:23,525 --> 00:05:44,700 [Hal Turing] So the imbalance isn't a bug you patch — it's a direct consequence of how attention's complexity scales, colliding with how real users send requests. Fixing it can't just mean tweaking the scheduler on top of the existing architecture — you'd have to rethink how attention and the experts talk to each other. So let's get into what they did. 12 00:05:44,700 --> 00:06:26,825 [Dr. Ada Shannon] That's exactly the redesign they went with, Hal. Instead of coupling attention and MoE math on the same chips the way every synchronous system does, ASAP physically splits them apart — that's what the paper calls disaggregated attention-expert serving. Attention lives on D times T devices, D data-parallel groups each split T ways for tensor parallelism, and the experts get an entirely separate pool of E devices. What bridges the two isn't a network call in the usual sense — every device on both sides carves out a chunk of memory that's visible to every other device in the system. Less like phones ringing each other, more like a wall of labeled mailboxes anyone can walk up to. That shared buffer is the whole trick: a decentralized hub, not a central switchboard. 13 00:06:26,825 --> 00:06:40,500 [Hal Turing] Okay, mailboxes I can picture — but mailboxes still need rules about who writes and who reads. What's the traditional way this data movement gets done, and what exactly are these four primitives replacing? 14 00:06:40,500 --> 00:07:39,575 [Dr. Ada Shannon] They replace what's called all-to-all collective communication — the standard blocking, handshake-based pattern where a device sending tokens out to the experts, or combining results back, has to negotiate directly with every recipient and wait for acknowledgment before it can move on. ASAP swaps that for four asynchronous primitives, paired by direction. async-dispatch-send and async-dispatch-recv move tokens from attention devices to the experts after each attention layer; async-combine-send and async-combine-recv bring results back. All four follow the same pattern: write, then flag. A sender writes straight into the target's shared buffer, flips a bit marking the slot full, and immediately returns to computing — no waiting for acknowledgment. The receiver polls those flags on its own schedule, and once everything it needs has landed, copies the data to private memory and clears the flag. One guardrail remains: if a sender tries to write into a slot whose flag is still set from the last round, it blocks until the receiver clears it. 15 00:07:39,575 --> 00:07:59,425 [Hal Turing] Oh wait, hold on — so it's basically fire-and-forget except for that one seatbelt clause. That's clever. But doesn't breaking one big batch into all these independent little transfers risk killing your compute efficiency? I'd assume MoE kernels want big fat batches, not a trickle of small ones. 16 00:07:59,425 --> 00:08:43,851 [Dr. Ada Shannon] Exactly the risk they flagged, and it's why four optimizations exist. Length-aware batching won't release a batch to the experts until it clears roughly two thousand tokens — the inflection point where MoE flips from memory-bound to compute-bound, so smaller batches would waste the hardware. Dual-batch interleaving pairs two batches per attention group and interleaves their layer-by-layer work, so devices never sit idle while one batch is off with the experts. Communication-computation overlapping is a triple-stream design — one stream for the math, two dedicated streams for async sends and receives, so communication never steals compute cycles. And the MoE Super Kernel makes the kernel 'layer-oblivious,' with pre-calculated address offsets for every layer's weights, so the CPU can dispatch kernels ahead of time instead of stalling to figure out which layer runs next. 17 00:08:43,851 --> 00:08:57,451 [Hal Turing] That's a lot of engineering stacked on top of the disaggregation itself. Did it actually translate into numbers, or is this one of those papers where the architecture diagram is prettier than the results table? 18 00:08:57,451 --> 00:09:58,026 [Dr. Ada Shannon] The numbers back it up. Under a five-second TTFT SLO, ASAP sustains 20 requests per second, where Default — their vLLM-like baseline — tops out at 6.8, and ChunkedPrefill maxes at 10.5. That's a 194% gain over Default, 90% over ChunkedPrefill. You can see exactly where it comes from: at RPS=4, they decomposed TTFT for short requests under Default and found synchronization waiting ate 55% of total latency, with another 30% lost to plain queuing — barely 15% was actual compute. ASAP cuts that non-kernel overhead by up to 80% for those same short requests. The communication layer is fast in absolute terms too — async-dispatch runs 4 to 5.8 times faster than synchronous point-to-point, under 0.1 milliseconds for 512 tokens, versus the 0.5 milliseconds reported by Liu, Tian, Wang and colleagues' 2025 EaaS system, though that's on different hardware. 19 00:09:58,026 --> 00:10:11,676 [Hal Turing] Okay, with four separate optimizations stacked together, I always wonder if one of them is secretly carrying the whole system while the others just round out a nice chart. Which one actually moved the needle most? 20 00:10:11,676 --> 00:10:55,301 [Dr. Ada Shannon] Dual-batch interleaving is the biggest single lever — it takes SLO-compliant throughput from 17.5 up to the full 20 RPS, a 14.3% gain, because it directly attacks attention-device idle time. Comm-compute overlapping is close behind at 12.4%, pushing from 17.8 to 20 RPS, and it matters more as load rises — at low RPS there's so much slack that hiding communication behind compute barely registers. The MoE Super Kernel is the smallest of the three at 6%, but it's a clean, almost arithmetic win: host-side kernel dispatch costs about 220 microseconds per layer, and DeepSeek V3.2 has 61 layers, so eliminating that dispatch stall saves roughly 13 milliseconds of TTFT per request. 21 00:10:55,301 --> 00:11:06,176 [Dr. Ada Shannon] But before we get too impressed with our own chart, Hal — all four of those percentages came from ASAP's team measuring ASAP against baselines ASAP's team built. 22 00:11:06,176 --> 00:11:31,676 [Hal Turing] That's exactly what nagged at me in the eval section. 'Default' and 'ChunkedPrefill' aren't vLLM and Sarathi — they're in-house reimplementations the paper itself calls 'similar to' those systems. And their own related-work section names MegaScale-Infer and Step-3 as the actual closest comparison — real attention-MoE disaggregation systems — and never runs a head-to-head against either one. 23 00:11:31,676 --> 00:12:12,776 [Dr. Ada Shannon] Guilty as charged, worth naming names. 'Default' mirrors vLLM — Woosuk Kwon and Ion Stoica's group at UC Berkeley, SOSP 2023. 'ChunkedPrefill' is Sarathi, Amey Agrawal and coauthors, Microsoft Research, 2023. MegaScale-Infer is Ruidong Zhu and the ByteDance team, SIGCOMM 2025; Step-3 is Bin Wang and coauthors, out of StepFun, 2025. Both get one dismissive sentence — 'relies on synchronous barriers,' 'targets decoding' — no benchmark run against either. And every baseline executes on the authors' own CloudMatrix384, so the 90-to-194 range is really ASAP against in-house code on home turf, not against the field. 24 00:12:12,776 --> 00:12:35,276 [Hal Turing] Oh wait, hold on — that connects to something in the abstract itself. It headlines '90% improvement,' but Section 5.2 reports 194% over Default and only 90% over ChunkedPrefill. They put the smaller number on the cover. Responsible caution, or does it quietly undersell how bad the weakest baseline looks? 25 00:12:35,276 --> 00:12:56,526 [Dr. Ada Shannon] Both, honestly, and that's what makes it useful rather than a gotcha. Leading with the harder-to-attack number is the more defensible choice — I'd trust that more than a paper opening with its biggest number. But it also means the abstract tells you nothing about which comparison it's from unless you go dig into Section 5.2. General rule: a headline percentage is a claim about one specific baseline, never a property of the system itself. 26 00:12:56,526 --> 00:13:23,176 [Hal Turing] Same instinct applies to something buried in the communication design. async-dispatch-send has a backpressure clause — if the target flag's already set, the write just blocks until the receiver clears it. That's a synchronization point sitting inside a system whose entire pitch is eliminating synchronization. Every result they report tops out at RPS=20. What happens at 21, or if a receiver just wedges? 27 00:13:23,176 --> 00:14:17,726 [Dr. Ada Shannon] There's nothing in the paper about that — no overload curve past the SLO ceiling, no fault injection. In a shared-buffer design, one receiver dying while holding a set flag can hang every sender targeting that buffer — the same straggler problem this whole system was built to kill, just relocated into the plumbing. Their own predecessor work, Guowei Liu and coauthors, 2026, actually coined 'DP Imbalance' — ASAP borrows the term but never says what fix that paper proposed. Same pattern with EaaS, Ziming Liu, Boyu Tian, and colleagues, 2025 — neither preprint even lists an institution, and ASAP only borrows EaaS's published latency number off a different hardware platform, GPU with IBGDA versus their own Ascend fabric, not a controlled same-hardware test. Funny footnote: one of ASAP's own authors, Ming Yan, also shows up on this year's Qwen3Guard report — common enough name that I won't swear it's the same person. 28 00:14:17,726 --> 00:14:43,126 [Hal Turing] Zooming out further — everything here is one model, DeepSeek-V3.2, one parallelism config, one supernode. The paper claims the async paradigm is 'platform-agnostic' and would work fine over commodity RDMA, and that it could extend to decode too, but neither claim gets an actual experiment. So practically, how much of this 90-to-194 range survives outside Huawei's own hardware? 29 00:14:43,126 --> 00:15:23,876 [Dr. Ada Shannon] Honestly, we don't know, and the paper admits as much if you read closely. The RDMA portability argument is entirely qualitative — hundreds of microseconds is still small next to the tens of milliseconds barriers cost, no data shown. Decode is explicitly flagged as a smaller win: decode attention scales linearly with sequence length instead of quadratically, so there's less imbalance to begin with, and per-request token counts are too small to saturate the MoE stage. What's genuinely new here isn't any single primitive — it's packaging four asynchronous mechanisms into one barrier-free pipeline that holds up at RPS=20. What's unproven is whether that holds on anyone else's hardware, model, or baseline code. 30 00:15:23,876 --> 00:15:42,051 [Hal Turing] So if I'm running MoE inference on a commodity GPU cluster, the honest takeaway isn't 'go implement this today' — it's 'barrier-free execution is a viable direction, worth an independent benchmark on your own stack before you trust the headline number.' 31 00:15:42,051 --> 00:16:01,451 [Dr. Ada Shannon] That's the right level of confidence. Architecturally this is a real contribution — dismantling global synchronization barriers is a legitimate idea, not a rebrand. Whether it travels beyond one vendor's supernode and one team's own baselines is the open question, and it's on someone outside Huawei to answer it. 32 00:16:01,451 --> 00:16:36,151 [Hal Turing] Good place to land, Ada. Bottom line for our listeners: ASAP shows that disaggregating attention and MoE onto separate hardware with genuinely non-blocking communication can meaningfully cut the synchronization stalls DP imbalance causes — that part's real and measured. What's still open is whether it generalizes past DeepSeek-V3.2, past CloudMatrix384's supernode fabric, and past baselines the authors built themselves. Worth watching for an independent replication. Thanks for listening, everyone — that's all for this one.