1 00:00:01,000 --> 00:00:35,975 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. And today we're digging into a paper called EventTensor: A Unified Abstraction for Compiling Dynamic Megakernel. First author Hongyi Jin, et al. — twenty-one authors total — out of Carnegie Mellon University, with co-authors from Shanghai Jiao Tong, NVIDIA, UC Berkeley, Princeton, Tsinghua, and Peking University. It hit arXiv April 21st, 2026, and it's published at MLSys 2026. 2 00:00:35,975 --> 00:00:57,225 [Hal Turing] So let's get concrete, because "events as first-class tensors" sounds nice on a slide, but I want to see it in code. The paper's got this split-K summation example in Figure 3 — summing across a row of a matrix. Walk me through it, Ada, because seeing the device function and graph function side by side is what makes this click. 3 00:00:57,225 --> 00:01:38,599 [Dr. Ada Shannon] You've got two device functions. partial_sum takes a tile coordinate i, j and writes a partial row-sum into B. final_sum takes coordinate i and sums B row i's partials into the final output C. Kernel-by-kernel, final_sum would wait for all of partial_sum to finish, even though C row i only actually needs B row i. Event Tensor lets you say that explicitly: you declare an ETensor of shape n with wait_count 4, partial_sum's out_edges map "ij->i" — every i,j tile notifies event i — and final_sum's in_edges are "i->i", so task i just waits on its own event. Row zero starts the second all four of its partial sums land, regardless of row fifty. 4 00:01:38,599 --> 00:01:53,049 [Hal Turing] That's encoding the true dependency instead of "wait for the whole kernel." But batch size changes every request in production. Does that mean regenerating this whole symbolic graph every time the shape changes? 5 00:01:53,049 --> 00:02:15,474 [Dr. Ada Shannon] No — that's Figure 4. The Event Tensor's shape itself can hold a symbolic variable, batch size B, whatever. So what you just described isn't compiled per shape, it's a template. At runtime it gets instantiated with whatever B actually is — a one-by-two event grid for batch size one, two-by-two for batch size two — no recompilation, no CUDA Graph recapture. That's the real break from graph-capture systems, which bake in one fixed sequence and choke the moment the shape moves. 6 00:02:15,474 --> 00:02:35,525 [Hal Turing] Shape dynamism's the easy case though — batch size is knowable before you start. MoE routing is the nasty one, since you don't know which expert gets which token until the router runs. How does something templated ahead of time handle a dependency that doesn't exist until runtime data shows up? 7 00:02:35,525 --> 00:02:51,700 [Dr. Ada Shannon] Two mechanisms. First, data-dependent event update — each expert's event has a counter, but instead of a fixed wait_count baked in at compile time, it's initialized at runtime straight from the topk tensor, the actual count of tokens the router assigned it— 8 00:02:51,700 --> 00:03:04,325 [Hal Turing] Oh wait wait wait — the synchronization primitive itself, the counter, is a runtime value? That's not just dynamic data flowing through a fixed graph, that's a dynamic dependency structure. 9 00:03:04,325 --> 00:03:36,375 [Dr. Ada Shannon] Exactly, and that's what conventional task graphs can't express. Second mechanism is data-dependent task triggering: exp_indptr, the prefix sum of tokens per expert, tells the scheduler exactly which range of GroupGEMM tiles to fire — expert i triggers tiles from exp_indptr of i to exp_indptr of i plus one. Both topk and exp_indptr are computed at runtime, and the kernel still compiles ahead of time, zero compilation overhead at inference — which matters once we get to the warmup numbers. 10 00:03:36,375 --> 00:03:49,175 [Hal Turing] So how does encoding all that turn into a runnable schedule? You've got static and dynamic scheduling transformations, and the GEMM-plus-Reduce-Scatter walkthrough is the concrete example. 11 00:03:49,175 --> 00:04:32,150 [Dr. Ada Shannon] Static scheduling precomputes a task queue per SM before launch, then lowers every Event Tensor edge into an explicit notify or wait call. Each Reduce-Scatter tile depends on two GEMM tiles, so its counter starts at two — SM0 finishes its tile, notifies, counter drops to one and SM0 spin-waits; SM1 finishes its tile, counter hits zero, and SM0's wait releases. Dynamic scheduling tracks the same dependency on-GPU instead: the moment a counter hits zero, it atomically pushes the ready task into a shared queue, and any idle SM pops it. There's also an early-push optimization — a task gets pushed the instant it's ready rather than waiting on the consumer to ask, so scheduling latency gets hidden off the critical path. 12 00:04:32,150 --> 00:04:42,475 [Hal Turing] Given dynamic handles both kinds of dynamism, why ever reach for static? Just default to dynamic and let the GPU sort out load balancing. 13 00:04:42,475 --> 00:05:13,150 [Dr. Ada Shannon] I actually disagree with you there, Hal. Dynamic isn't free — every push and pop is an atomic op against a centralized queue in global memory, and the paper flags contention as a real risk at scale. Table 3's TP=4 numbers back this up: on regular, predictable dense layers, static wins because there's zero runtime queue overhead, the schedule's already sitting in memory. Dynamic only pulls ahead on Table 2's irregular workloads, MoE, where completion order genuinely isn't knowable in advance. 14 00:05:13,150 --> 00:05:29,700 [Hal Turing] Fair — I was treating dynamic as strictly dominant because it's more general, but general isn't free. So it's not "dynamic is better," it's matching the scheduler to whether the irregularity is real unpredictability or just a large-but-fixed shape. 15 00:05:29,700 --> 00:05:38,500 [Dr. Ada Shannon] Right — MoE's irregularity is real. TP=4 dense GEMMs are just big, not unpredictable. 16 00:05:38,500 --> 00:05:43,000 [Hal Turing] Alright, give me the scoreboard — what did all this actually buy them? 17 00:05:43,000 --> 00:06:25,825 [Dr. Ada Shannon] On the fused kernels, up to 1.40x over cuBLAS-plus-NCCL on both GEMM-plus-Reduce-Scatter and All-Gather-plus-GEMM. On the standalone MoE layer — Qwen3-30B-A3B, 128 experts, top-k 8 — up to 1.23x against Triton and FlashInfer at 1024 tokens. End-to-end, at batch size one on that same model, they're 1.48x over vLLM and 1.20x over SGLang. And the number I find most operationally interesting is warmup: 35 seconds to get ETC's engine ready versus 123 for vLLM and 583 for SGLang — the ahead-of-time symbolic-shape compilation paying off, no recapture loop, no shape-specific recompilation. 18 00:06:25,825 --> 00:07:08,500 [Hal Turing] Okay, before we wrap, I want to poke at that 1.23x MoE number, because when I actually looked at how it was measured, it's narrower than the headline makes it sound. That's a single fused MoE layer, in isolation, run as a microbenchmark — not the full decoding pipeline. The paper's end-to-end number is the more honest measurement of what a user actually gets. But that layer-level 1.23x gain — has it been shown to generalize to shared-expert MoE, sigmoid gating, or something like DeepSeek's 256-expert routing? Because those have very different load-balancing profiles than top-k 8 out of 128. 19 00:07:08,500 --> 00:07:48,775 [Dr. Ada Shannon] No, and that's a real gap, not a nitpick. The dynamic scheduler's whole pitch is adaptive load balancing for irregular routing, and irregularity scales with expert count and gating shape. Sigmoid gating in particular doesn't give you the clean top-k selection this paper's data-dependent event update and task triggering mechanism was built around — the triggering logic assumes you can compute exp_indptr cleanly from a route decision. A shared-expert design, where every token also always hits one dense expert, changes the dependency shape entirely. None of that is tested here. So the 1.23x is a real result for one configuration, and everything past that is extrapolation the paper doesn't earn. 20 00:07:48,775 --> 00:08:13,125 [Hal Turing] Right, and that same caution applies even harder to the TP=4 result. Figure 14, the right panel — ETC only lands between 0.99x and 1.06x against vLLM there, and it actually trails SGLang. That's the regime where you'd most want the fine-grained scheduling to pay off, distributed dense layers, and it basically doesn't. 21 00:08:13,125 --> 00:08:45,400 [Dr. Ada Shannon] The paper's own explanation is that it's engineering, not the abstraction — compiler-generated GEMM tiles that aren't as tuned as cuBLAS yet, plus more CPU-side overhead in their serving engine than SGLang's scheduler. And I think that's plausible. Event Tensor's job is expressing and scheduling dependencies; it was never claiming to out-tune a decade of cuBLAS kernel engineering. Those are genuinely separable problems — you can bolt a better GEMM codegen backend onto the same Event Tensor graph later without touching the abstraction at all. 22 00:08:45,400 --> 00:09:11,350 [Hal Turing] Hold on, I actually disagree with you there, Ada — that's letting them have it both ways. If the headline claim is that fine-grained GPU-side scheduling beats coarse kernel boundaries, and the one head-to-head test against a mature, heavily-tuned serving engine comes out flat or behind, you can't just wave that off as 'someone else's tuning problem.' The abstraction's value proposition is supposed to show up precisely in a case like this. 23 00:09:11,350 --> 00:09:40,176 [Dr. Ada Shannon] I get the instinct, but I don't think it holds up mechanically. The scheduling win and the GEMM-tuning win are orthogonal axes — you can have a perfectly-scheduled megakernel built on mediocre tiles and still lose to a mediocre schedule built on excellent tiles. The honest read is we don't yet know how much of TP=4's result is Event Tensor's ceiling versus a fixable codegen gap, because nobody's run that ablation. That's not a defense of the paper, it's exactly the missing experiment — and it rhymes with the bigger omission here. 24 00:09:40,176 --> 00:10:08,151 [Hal Turing] Which is Mirage and Hazy Research. The paper positions Event Tensor as generalizing the manual scheduling in Mirage Persistent Kernel — Cheng, Zhang, Zhou and a long author list overlapping heavily with this very paper, out of CMU and collaborators, 2025 — and the hand-crafted Hazy Research Llama-1B megakernel from Spector, Juravsky, Sul and colleagues at Stanford, also 2025. Both get cited constantly. Neither gets benchmarked. 25 00:10:08,151 --> 00:10:48,501 [Dr. Ada Shannon] And the excuse given — that existing megakernel frameworks 'cannot be fairly compared under dynamic-shape or data-dependent workloads' — doesn't cover the static, single-shape cases, where a fair comparison should've been possible. That's the omission that actually matters for judging incremental contribution. Worth naming CuSync too, from Jangda, Maleki, Mehri Dehnavi, Musuvathi and Saarikivi at Microsoft Research, 2024 — it already does fine-grained dependent-kernel synchronization across streams. Event Tensor's real difference is making events first-class tensors in the compiler IR instead of co-scheduling kernels on separate streams, but that distinction deserved a number, not just a paragraph in related work. 26 00:10:48,501 --> 00:11:13,701 [Hal Turing] Setting the benchmarking gaps aside, the practical stuff still matters regardless of who wins TP=4. The paper says ETC is already folded into a major open-source system, and that warmup-time edge is an operational win independent of the throughput debate. For autoscaling and cold-start-sensitive serving, that's real money saved even if the raw speedup arguments stay unsettled. 27 00:11:13,701 --> 00:11:51,701 [Dr. Ada Shannon] Agreed, and it's the more defensible claim in the paper honestly. On where this goes next, the authors flag wanting to auto-generate Event Tensor task graphs from standard computation graphs, cutting the manual annotation burden — right now someone still has to hand-write the in_edges and out_edges wiring we walked through earlier. And I'd note the DSL-agnostic claim, that this slots into Triton or CuteDSL without conflict, is asserted, not demonstrated — everything here runs on their own TVM-based DSL. That's the scope-versus-claims gap: broad generality asserted, narrow generality tested. 28 00:11:51,701 --> 00:12:23,051 [Hal Turing] So here's where I land. The real contribution isn't a uniformly dominant speedup — it's compiler-level generality: one abstraction, both scheduling strategies, both flavors of dynamism, ahead-of-time for any shape. The MoE microbenchmark and the TP=4 result are honest limits, not fatal flaws, but they're limits worth remembering next time someone quotes that 1.23x without the fine print. That's it for this one — thanks for listening, we'll catch you next time.