1 00:00:01,000 --> 00:00:49,622 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called "When Does Disaggregation Pay? Simulating Prefill-Decode-Attention-FFN Specialization for Agentic LLM Inference." It's got eleven authors, led by Przemyslaw Forys et al., out of Imperial College London and the University of Cambridge, submitted to arXiv on August 4th, 2026. And Ada, the number that jumped out at me right away is buried in the abstract: up to a 2.06x throughput gain from splitting inference into four independently specialized hardware pools. But then in the very next sentence they basically say 'this only works sometimes.' That's an unusual thing for a systems paper to admit up front. 2 00:00:49,622 --> 00:01:27,935 [Dr. Ada Shannon] Right, and that's exactly why this one's worth a full episode instead of a footnote. Everybody in serving infrastructure already agrees prefill and decode want different hardware — that's old news at this point. What this group is asking is sharper: if you keep cutting, splitting attention from FFN too, so you get four independently provisioned stages instead of two, does that extra complexity actually buy you anything, or is it just more knobs to misconfigure? And their answer is genuinely conditional — it depends on your workload's prefill-to-output ratio and on whether your hardware even lets you tune compute and memory independently. That second part turns out to matter a lot more than I expected. 3 00:01:27,935 --> 00:01:45,582 [Hal Turing] Let's set the stage, because I think a lot of listeners hear 'agentic workloads' and nod along without picturing what's actually different. What's actually happening under the hood when a model is, say, browsing the web or calling fifteen tools in a row instead of just answering one chat message? 4 00:01:45,582 --> 00:02:22,920 [Dr. Ada Shannon] So the paper cites some numbers here that are pretty stark. A standard chatbot turn averages around 1,200 tokens of context. An agentic computer-use benchmark called OSWorld averages 38,000 tokens and can spike past 100,000. That's because every tool call, every screenshot, every intermediate reasoning step gets stuffed back into the prompt on the next turn. You're re-prefilling an ever-growing context almost every turn, not just once at the start of a conversation. And that changes the shape of the workload enough that a single homogeneous bank of GPUs, sized for chatbot traffic, starts to buckle. 5 00:02:22,920 --> 00:02:28,771 [Hal Turing] Okay, so walk me through prefill and decode themselves, because I think that's the foundation everything else builds on. 6 00:02:28,771 --> 00:03:10,474 [Dr. Ada Shannon] Prefill is when the model reads your entire input prompt at once and produces the first output token. Because it's processing many tokens in parallel, it's compute-bound — you're doing a ton of matrix multiplication and GPUs are happy doing that. Decode is the opposite: one token at a time, autoregressively, and at each step you have to read the entire key-value cache built up so far just to produce one new token. As that cache grows with context length, decode becomes memory-bandwidth-bound — you're not short on FLOPs, you're short on the ability to move bytes off memory fast enough. Same chip, two completely different bottlenecks, which is a genuinely awkward thing to size hardware for. 7 00:03:10,474 --> 00:03:21,806 [Hal Turing] Wait, hold on — sorry to cut you off, but that's the part that always gets me. If it's the exact same physical GPU sitting there for both phases, isn't it always going to be wrong for one of them? 8 00:03:21,806 --> 00:03:59,422 [Dr. Ada Shannon] Pretty much, yeah, and that mismatch is precisely why prefill-decode disaggregation exists as a real, shipping idea and not just an academic thought experiment. DistServe, out of Peking University, OSDI 2024, and Splitwise, from Microsoft Azure, ISCA 2024, both put prefill and decode on physically separate device pools so each pool can be sized and scheduled for its own bottleneck instead of compromising. You pay a cost — the KV cache has to be shipped over the network from the prefill machine to the decode machine — but for latency-sensitive serving at scale, both papers showed it's worth it. 9 00:03:59,422 --> 00:04:04,902 [Hal Turing] And this paper wants to go one cut deeper than that split, right? Into attention and FFN separately? 10 00:04:04,902 --> 00:05:03,881 [Dr. Ada Shannon] Exactly, and that's not a hypothetical either — StepFun did this in production first, in their Step-3 system, splitting attention and FFN onto separate hardware because the two sublayers batch completely differently. Attention reads a distinct KV cache per request, so a bigger batch just adds memory traffic with no reuse. FFN applies the same weights to every token, so a bigger batch means more reuse and higher utilization. Put them on the same device and you're stuck compromising again, just one level down. So this paper's contribution, PDAF, takes it all the way: prefill-attention, prefill-FFN, decode-attention, and decode-FFN each get their own independently specialized pool. Four stages, four different hardware profiles, and — per that headline number — up to 2.06x throughput when it works. The catch is 'when it works' is doing a lot of heavy lifting, and that's where we're headed next. 11 00:05:03,881 --> 00:05:38,385 [Dr. Ada Shannon] To actually test whether that four-way split is worth it, they built a simulator called HeteroPanacea, and the thing I like is they didn't just trust it — they validated it against a real 8-node cluster of NVIDIA B200s. They compared their roofline predictions for compute throughput against measured GPU kernels, and they compared their power model against actual datasheet TDP. Both come out close in their Figures 2 and 3. So this isn't a napkin-math simulator pulled out of thin air — the underlying compute and power models are anchored to hardware that exists today, even though the exotic custom chips it searches over don't. 12 00:05:38,385 --> 00:05:47,302 [Hal Turing] Okay, and how does it actually decide what happens over time — like, does it model requests queuing up realistically, or is it more of an idealized average? 13 00:05:47,302 --> 00:06:25,383 [Dr. Ada Shannon] It's event-driven, which is the right level of fidelity for this kind of design-space search. Prefill requests get accumulated greedily until a timeout flushes the batch, and decode admits requests FIFO, bounded by how much KV-cache capacity is actually free. But — and this matters for how much to trust the headline numbers — it doesn't model queuing delay or kernel-launch overhead below that per-stage granularity. Every stage's time is just the roofline bound: max of compute time or memory time. So it's fast and it scales to eight models and nine workload ratios, but it's an idealized greedy scheduler, not a faithfully reproduced production stack. 14 00:06:25,383 --> 00:06:34,346 [Hal Turing] That's a good segue into the hardware side, because I was surprised reading this — they didn't just simulate existing chips, they invented a whole synthetic design space. 15 00:06:34,346 --> 00:07:22,968 [Dr. Ada Shannon] Yeah, this is Table II, and it's the part that made the paper click for me. On a real GPU, compute and memory bandwidth come bundled — an H100 is an H100, you can't buy the compute of an H100 with the memory bandwidth of something else. Here, they let compute range from 25 up to 20,000 teraflops completely independently of memory, and memory itself is drawn from seven technology classes — SRAM, HBM, DDR, LPDDR, GDDR — each with its own capacity-bandwidth tradeoff. So a stage can get, say, massive compute paired with tiny, cheap memory, or the reverse. That decoupling is the whole trick, and it's exactly what real accelerators can't do yet, though it's basically the bet NVIDIA's Vera Rubin platform is making by pairing Rubin GPUs with Groq LPUs. 16 00:07:22,968 --> 00:07:28,680 [Hal Turing] So walk me through the actual payoff number, because 2.06x is the headline everyone's going to quote. 17 00:07:28,680 --> 00:08:13,448 [Dr. Ada Shannon] Right, and the crucial word is conditional. On the custom NPU space, PDAF averaged across all eight models hits that 2.06x peak, but only once the workload becomes prefill-heavy. Figure 5 sweeps nine prefill-to-output ratios and the transition is abrupt — PDAF is sitting at 0.48x, actually worse than doing nothing, at a balanced ratio, and then jumps to 2.10x just one decade later. And it's not monotonic past that: it plateaus around 1.5 to 1.8x through a very high ratio and then actually falls back to 1.27x at the extreme end, because once prefill totally dominates, the decode hardware you provisioned separately just sits idle. Agentic workloads happen to land right in that sweet middle band. 18 00:08:13,448 --> 00:08:22,597 [Hal Turing] Wait, hold on — sorry, jumping in — so at a normal chatbot-style balanced ratio, splitting into four stages actually makes things worse than just leaving it alone? 19 00:08:22,597 --> 00:09:15,678 [Dr. Ada Shannon] Exactly, below parity, because you're paying the coordination overhead of four separate pools without enough asymmetry between the stages to exploit. And that crossover point looks completely different once you swap the synthetic NPUs for real GPU catalog hardware, which is Section VI-C. On GPUs, plain PD disaggregation clears parity almost immediately, even at very low prefill ratios, winning for all eight models. But PDAF barely helps there — and Table V explains exactly why: when they look at the actual per-stage hardware the search picks, 31 out of 32 PDAF stage assignments across every model are just... an H100. One outlier is an A100. Because GPUs tie compute and bandwidth together, there's no independent knob to turn, so splitting attention from FFN buys almost nothing beyond what splitting prefill from decode already got you. 20 00:09:15,678 --> 00:09:26,870 [Hal Turing] That's a pretty stark way of showing PDAF's benefit is really a hardware-availability story, not just an algorithm story. What about the quantization piece — I know they ran that on a smaller model. 21 00:09:26,870 --> 00:10:14,749 [Dr. Ada Shannon] Qwen3.5-32B, and they tested two very different workload shapes: BFCL, which is heavily prefill-dominated at over eleven thousand tokens of context, and a GSM8K subset that's much shorter. Starting from an 8-bit baseline everywhere, they dropped exactly one stage to 4 bits at a time. Push either FFN stage down and GSM8K accuracy craters — down to 15 or 47 percent from a 75 percent baseline — while BFCL barely moves. Flip it and push either attention stage down instead, and BFCL collapses to around 11 or 12 percent while GSM8K is basically untouched. So which stage tolerates low precision is a property of the workload's shape, not some fixed property of the model. 22 00:10:14,749 --> 00:10:21,530 [Hal Turing] And then there was that ablation section, which I thought was the most useful part for actually predicting when to bother with any of this. 23 00:10:21,530 --> 00:11:16,329 [Dr. Ada Shannon] It's the payoff of the whole paper, honestly. They swept five architecture factors independently. Decode-attention's KV traffic — controlled by the latent rank in MLA-style attention — moves PDAF's benefit monotonically: widen that rank and the gain climbs steadily upward across all three models they tested, because a wider latent vector makes attention more clearly KV-bandwidth-bound and more clearly distinct from FFN. Growing total expert count while holding active experts fixed is nearly free — flat within a few percent, because that only adds capacity pressure, not compute or bandwidth pressure. But growing active-expert count, real compute per token, is the one lever that collapses the benefit outright, one model loses over 70 percent of its gain. And precision cuts differently per model depending on whether narrowing it preserves or erases the gap between attention and FFN. 24 00:11:16,329 --> 00:11:53,620 [Dr. Ada Shannon] Right, the whole time, from 128 up to 4096. Physically that number is the width of the compressed KV vector MLA caches per token, so a wider rank just means more bytes crossing the wire on every decode step, and that's exactly the traffic decode-attention's dedicated pool exists to soak up. At the narrow end it's the opposite story: decode-attention is basically weight-read dominated, indistinguishable from decode-FFN, so the four-way split has nothing left to exploit and actually falls below plain non-disaggregated serving for two of the three models they tested. 25 00:11:53,620 --> 00:12:32,119 [Hal Turing] Okay, so here's where I want to push on the paper itself, Ada, not just the results. They validate the roofline compute, power, and communication pieces against a real 8xB200 node — Tables III and IV, the figures we walked through — but the scheduler that actually produces the 2.06x number and that crossover decade, greedy prefill accumulation, FIFO decode, zero queuing overhead below the stage level, none of that gets checked against a real serving stack. How much should we actually trust the magnitude here versus just the direction? 26 00:12:32,119 --> 00:13:24,967 [Dr. Ada Shannon] The paper says as much itself in the limitations, to its credit. And there's a useful contrast: DistServe, from Zhong, Liu, Chen and colleagues out of Peking University and collaborators, published at OSDI in 2024, validated its prefill-decode gains on a live serving stack, not a simulator. Same with Splitwise, Patel, Choukse, Zhang and coauthors out of Microsoft Azure and Microsoft Research, ISCA 2024 — real phase-splitting deployment. Those papers earned their throughput numbers the hard way. HeteroPanacea's qualitative claim, disaggregate when the workload is prefill-heavy, rhymes with what both of those found. But I'd treat 2.06x and the exact I/O=10 crossover as an upper-bound estimate from an unvalidated scheduler layered on top of a validated roofline, not a number you'd write into a procurement spec. 27 00:13:24,967 --> 00:13:44,426 [Hal Turing] And that unvalidated-scheduler problem gets worse once you add in the custom NPU space, right — the independent compute-memory scaling from Table II that no real chip actually has. So is PDAF's headline advantage a real hardware opportunity, or an artifact of simulating chips that don't exist yet? 28 00:13:44,426 --> 00:14:22,878 [Dr. Ada Shannon] Bit of both, honestly. It's a real opportunity in the sense that Vera Rubin is explicitly aiming at that independent-tunability regime, so this isn't pure fantasy hardware — but it's contingent on that platform actually shipping with the flexibility the paper assumes. For the record, that's Step-3's paper title: 'model-system co-design for cost-effective decoding,' 2025. And here's the humbling part: this paper's own AF-only numbers never clear parity, best case 0.96x on GPUs — even though AF is the piece with real-world precedent, and PDAF, the untested extension, is what actually wins. 29 00:14:22,878 --> 00:14:54,410 [Hal Turing] Wait, hold on — sorry, jumping in — but that's a real problem for anyone actually trying to act on this. If PDAF's benefit swings from 0.93x to 1.85x on GPT-OSS just from moving kv_lora_rank, or collapses 73% on DeepSeek-V4-Flash from active-expert count, those are architecture choices a model vendor could flip in their next release. You're talking about committing years of hardware capital to a ratio that might not survive the next model generation. 30 00:14:54,410 --> 00:15:28,265 [Dr. Ada Shannon] That's the real tension, yes. Though there's a silver lining buried in that same ablation: the paper's finding is basically that PDAF's payoff is predictable straight from a model's config file — decode-attention's KV width and active FFN compute — without running the full hardware search at all. So an operator isn't stuck re-simulating everything every time a new model drops; they could screen candidate models against that predictor first and only commit hardware once several generations in a row land in the same regime, rather than betting on one snapshot. 31 00:15:28,265 --> 00:15:38,621 [Hal Turing] What's actually novel here versus incremental, then? Because PD disaggregation is old news, AF comes from Step-3, and PDAF is really just composing the two. 32 00:15:38,621 --> 00:16:37,228 [Dr. Ada Shannon] The composition and the simulator itself are the contribution, not the individual splits. Nobody had built a tool that lets you jointly search hardware, parallelism, and quantization across all four stages at once — that's genuinely useful infrastructure, and I'd put it ahead of most of the comparison table. Where the abstract oversells is calling this guidance for 'next-generation AI infrastructure' broadly. The quantization study, the part meant to connect precision choice to hardware choice, only ran on one model, Qwen3.5-32B, and two tasks, while every model in the throughput experiments is MoE — so those two searches never actually touched the same model. And the workload model has no notion of prefix caching, which real agentic serving increasingly leans on — Mooncake, Qin and colleagues, FAST 2025, built exactly that KV-cache reuse layer HeteroPanacea's request model ignores. That could shrink the prefill-heavy regime where PDAF wins in practice. 33 00:16:37,228 --> 00:16:52,600 [Hal Turing] So the honest takeaway: trust the direction, treat 2.06x as a ceiling, and check your own model's KV-cache width before you believe any of it applies to you. That's a good place to leave this one. Thanks for listening, everyone — we'll catch you next time.