1 00:00:01,000 --> 00:00:53,198 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper is 'Scaling Inference Prefill with High-Radix Photonic Interconnects,' by Arulselvan Madhavan et al. — four co-authors total, all out of Lightmatter — accepted at the 33rd IEEE Symposium on High-Performance Interconnects, HotI 2026, posted to arXiv on September 1st this year. Flagging this immediately: Lightmatter makes photonic interconnects, so this is a vendor benchmarking its own bet against copper. Keep that in mind. But here's the number that stopped me scrolling: buried in the results is a claimed 8.5x cut in prefill latency for long-context inference. Eight and a half times, Ada. Gut read, before we've even opened the methodology? 2 00:00:53,198 --> 00:01:31,743 [Dr. Ada Shannon] My gut read: show me the workload before I trust the multiplier. But the underlying question is legitimate — does swapping copper scale-up interconnects for 3D-integrated photonics meaningfully cut prefill latency and time-to-first-token for Mixture-of-Experts models, across short, medium, and long context? That's not academic anymore. Inference is now eighty to ninety percent of all AI compute cycles, and agentic coding workflows are pushing median prompt lengths to roughly ninety-six thousand tokens, with nearly half of requests topping 128K. Prefill — the entire subject of this paper — has quietly become the dominant cost center in production. 3 00:01:31,743 --> 00:02:05,087 [Hal Turing] Let's ground that for anyone newer to the space. Prefill is the model chewing through your entire input — system prompt, history, pasted documents — all at once, to build the initial key-value cache. It's compute-bound: limited by raw FLOPS. Time-to-first-token is basically how long that takes. Then decode kicks in, generating your response one token at a time, and that phase is bound by memory bandwidth instead, because every step has to pull the full model weights out of HBM. Same model, two completely different bottlenecks. 4 00:02:05,087 --> 00:02:42,053 [Dr. Ada Shannon] That memory-bandwidth story for decode is covered well in Pope, Douglas, et al., 'Efficiently Scaling Transformer Inference,' out of Google, MLSys 2023, if anyone wants the deep dive. Prefill has its own headache here, specific to Mixture-of-Experts. Instead of every token running through one dense network, a router sends each token to a handful of specialist 'expert' sub-networks out of potentially hundreds or thousands. Great for parameter efficiency, but it means tokens are constantly shuffled across devices in what's called an all-to-all — a big cross-device chatter problem stacked on top of the raw compute. 5 00:02:42,053 --> 00:02:54,035 [Hal Turing] So where does the interconnect physically come in? 'Scale-up' and 'scale-out' get thrown around constantly — what's actually different, and why would a company care so much about something called 'radix'? 6 00:02:54,035 --> 00:03:32,765 [Dr. Ada Shannon] Radix is just how many links a device fans out to. Scale-up is the tightly-coupled, ultra-high-bandwidth fabric inside a rack — think NVLink. Scale-out is the slower fabric connecting racks together, like InfiniBand. Every chip hits the same wall regardless of design cleverness: I/O bandwidth is capped by the 'shoreline,' literally the die's physical perimeter, since that's the only place wires can exit. Copper adds a reach problem on top of that — at 224 gigabits per second per lane, passive copper is only reliable out to about one meter. That's the whole reason scale-up pods get stuck at a single rack. 7 00:03:32,765 --> 00:03:48,787 [Hal Turing] A one-meter leash forcing everyone into multi-rack scale-out headaches. So this is where Lightmatter's actual product shows up, right? Walk me through what 3D-integrated photonics even means — 'stacking optics on a chip' sounds simple until you think about it. 8 00:03:48,787 --> 00:04:10,196 [Dr. Ada Shannon] So instead of routing every signal out through shoreline SerDes, their Passage platform stacks the electronic host chip directly on top of a photonic circuit — optical modulators maybe a hundred microns under the electrical die. That decouples I/O from the die's edge completely, so light can exit from anywhere across the whole chip surface instead of being squeezed through the perimeter— 9 00:04:10,196 --> 00:04:22,828 [Hal Turing] Oh wait, wait, hold on — so it's not that fiber is inherently faster than copper, it's that you're no longer rationed by how much of the edge you can wire up? That's a real-estate problem, not a speed problem. 10 00:04:22,828 --> 00:05:02,580 [Dr. Ada Shannon] Exactly right — a radix and bandwidth-density story, not a different signaling medium. The specs back it up: over 64 terabits per second bidirectional per GPU, versus 14.4 terabits on current NVLink-class Blackwell hardware. Fiber pitch is tighter too — about 127 microns versus 254 for copper — and one bidirectional fiber replaces four copper wires, roughly an 8x radix jump. Caveat worth flagging now: the paper only cites Passage's published spec sheet for these figures. They don't directly simulate Lightmatter's own optical hardware. That distinction matters for everything downstream. 11 00:05:02,580 --> 00:05:22,689 [Hal Turing] Fair, but even granting the skepticism, the physics constraint itself is real no matter who's telling us — shoreline is shoreline, one meter is one meter. So when they report gains ranging from about 2.1x up to 8.5x depending on context length and configuration, doesn't that at least pass the sniff test? 12 00:05:22,689 --> 00:05:52,921 [Dr. Ada Shannon] I actually disagree with you there, Hal — you're collapsing two different questions. That the shoreline constraint is real, sure, nobody's disputing the physics. Whether their specific 2.1-to-8.5x range reflects a real deployment is a separate claim entirely, and it comes from a simulator Lightmatter itself wrote and controls. A genuine physical constraint doesn't automatically validate a specific number from an interested party's own model. Different confidence levels — we shouldn't blur them together this early. 13 00:05:52,921 --> 00:06:08,246 [Hal Turing] Fair enough — worth pulling apart once we see how they actually built this simulator. For now: 2.1x at the low end, up toward 8.5x at the extreme, across three model sizes and four hardware generations. 14 00:06:08,246 --> 00:06:30,816 [Dr. Ada Shannon] And that spread is itself informative — it tells you the benefit isn't uniform, it depends on whether a given regime is compute-bound or communication-bound. That's the crux of this paper: mapping exactly when communication becomes the bottleneck for MoE prefill. To know whether their numbers hold up, we need to see how this thing was actually built — because none of it ran on real photonic hardware. 15 00:06:30,816 --> 00:06:50,042 [Hal Turing] That actually ties into something I want nailed down before we go further, Ada — every single one of these ratios comes out of a simulator these four authors built themselves. So what is this thing, mechanically? Because "we built a cost model" can mean a spreadsheet or it can mean something much more serious. 16 00:06:50,042 --> 00:07:39,733 [Dr. Ada Shannon] It's the latter, mechanically speaking. They maintain a fork of Google's XLA compiler and use it to walk the MLIR representation of a production-style model partition — tensor-parallel, expert-parallel, context-parallel, whatever mesh the model actually uses — and at each op the compiler emits, they capture an estimated compute time and an estimated communication time. So the collectives aren't hand-waved, they're the actual all-to-all and all-reduce calls the compiler would generate in a real deployment. But here's the load-bearing caveat: every number in this paper, every multiplier, comes from that cost model. There's no real photonic silicon and no matched electrical cluster actually running these workloads to check the estimates against. It's internally consistent, not independently verified. 17 00:07:39,733 --> 00:07:43,820 [Hal Turing] Right — so what are they actually running through it? What's in Table Two? 18 00:07:43,820 --> 00:08:36,065 [Dr. Ada Shannon] Three MoE tiers. Mini at 21 billion active parameters, R1 at 42 billion, aligned to DeepSeek-R1's architecture, and Next at 201 billion, meant to represent a next-generation frontier system. All three use multi-head latent attention for KV compression, and they test both FP4 and FP8 quantization. The wrinkle is the R1 variant isn't actually DeepSeek-R1 — they adjusted the embedding dimension to 8064 instead of 7168, bumped routed experts from 256 to 288, and expanded the vocabulary, purely so the model divides cleanly across device counts of 8, 16, 24, 48, and 72. End result is a model about 14% larger than the real DeepSeek-R1, same communication topology, different parameter count. 19 00:08:36,065 --> 00:08:57,984 [Hal Turing] So when people see 'DeepSeek-R1' in the headline number later, that's actually this divisibility-adjusted stand-in, not the model Chinese labs actually shipped. Good to flag. Now — hardware. They model four platforms, B200, B300, Rubin, and something called R4. What's R4? 20 00:08:57,984 --> 00:09:11,684 [Dr. Ada Shannon] R4 is where I want you to slow down, because it's doing a lot of work later. B200, B300, and Rubin are real roadmap platforms with public specs. R4 is a speculative quad-die configuration the authors inferred by extrapolating from — 21 00:09:11,684 --> 00:09:15,864 [Hal Turing] Wait, hold on — inferred from what, exactly? That's not a shipping part. 22 00:09:15,864 --> 00:09:46,282 [Dr. Ada Shannon] Correct, it isn't. It's extrapolated from a paywalled SemiAnalysis newsletter report describing a proposed dual-die Vera Rubin GPU, and R4 projects that forward into a hypothetical quad-die version with up to 576 GPUs in a scale-up pod. So we're two layers removed from anything that exists: a paywalled industry report about a chip that isn't shipped, extrapolated into a chip that doesn't exist at all, then run through a cost model nobody outside Lightmatter can audit. 23 00:09:46,282 --> 00:09:56,220 [Hal Turing] Noted, and we'll come back to why that matters for the headline number. Walk me through how they actually generate the multipliers — I saw references to a device sweep and a batch sweep. 24 00:09:56,220 --> 00:10:24,177 [Dr. Ada Shannon] Device sweep fixes total tokens and batch size per context length, then varies the GPU count and measures overlapped prefill latency, electrical versus optical, at every point. Batch sweep does the reverse — fixes device count, varies batch size. And for both, they don't report a typical point or an average — they report the maximum multiplier they find anywhere in the sweep. That's an important methodological choice to flag now, before we get to the numbers. 25 00:10:24,177 --> 00:10:27,195 [Hal Turing] Okay, so what do those maximum multipliers actually look like? 26 00:10:27,195 --> 00:11:06,298 [Dr. Ada Shannon] At 1K to 8K tokens, high-batch regimes, 2.1 to 2.9x. At 128K, once you cross from a single rack into scale-out — 288 GPUs and beyond — rack-limited platforms like B200 and Rubin hit as much as 5.8x, because that's exactly where electrical systems pay the cross-rack tax. At 1M tokens, production platforms land at 2.2 to 4.5x. And the R4 configuration, at 1152 devices, hits 8.5x — the single largest number in the whole table, and it's sitting on the platform we just established doesn't exist yet. 27 00:11:06,298 --> 00:11:16,422 [Hal Turing] Right. So separate from that bulk-prefill sweep, they also ran something closer to real serving conditions — a discrete-event simulation. What did that show? 28 00:11:16,422 --> 00:11:53,527 [Dr. Ada Shannon] Six 8-GPU prefill workers, deliberately paired with just one 24-GPU decode worker, Poisson arrivals at 250 requests per second. Optical scale-up cuts p99 TTFT by 12 to 20% — faster dispatch, faster queue drain. But because prefill now feeds the single decode worker faster and in tighter bursts, that worker's steady-state batch grows, and p99 time-per-output-token gets worse — 77 to 110% worse. Net effect: end-to-end p99 latency comes out roughly neutral, within about 5% either way. 29 00:11:53,527 --> 00:12:05,741 [Hal Turing] But TTFT is the metric everyone actually optimizes against for interactivity — so even with TPOT eating into it, doesn't that still read as a genuine win for the user-facing number that matters most? 30 00:12:05,741 --> 00:12:34,998 [Dr. Ada Shannon] I actually disagree with you there, Hal. If end-to-end latency is flat, the user isn't experiencing a win — they're experiencing a shifted bottleneck. This is precisely the failure mode Zhong and colleagues warned about in DistServe, disaggregating prefill and decoding for goodput-optimized LLM serving, out of Peking University and UC San Diego, published at OSDI 2024 — prefill and decode capacity have to be co-sized, or gains on one side just create pressure on the other. 31 00:12:34,998 --> 00:12:50,416 [Hal Turing] Fair — though isn't a single 24-GPU decode worker against six prefill workers a pretty extreme, almost worst-case ratio to begin with? That feels like it's engineered to expose the tradeoff rather than reflect how anyone would actually provision this. 32 00:12:50,416 --> 00:13:19,905 [Dr. Ada Shannon] That's a fair pushback, and I'll meet you partway — it is an isolating choice, deliberately so, to make the interaction visible. But it's also not far from the throughput-latency tension Agrawal and colleagues mapped in Sarathi-Serve, out of Georgia Tech and Microsoft Research India, OSDI 2024. Under-provisioned decode is a realistic failure mode, not a strawman. The authors even say as much themselves — this shows headroom, not a sizing rule. 33 00:13:19,905 --> 00:13:43,775 [Dr. Ada Shannon] Underneath the noise, yeah — Sarathi-Serve's whole contribution was chunking prefill so it doesn't starve decode's scheduling slots, which is a software answer to the exact tension this paper's DES experiment stumbles into with hardware. That's the connection I want to land before we move on: the prefill multipliers in Table Four aren't a deployment win by themselves — they're headroom, contingent entirely on decode keeping pace. 34 00:13:43,775 --> 00:14:21,949 [Hal Turing] Okay, headroom, not a finished product — I can hold that. But I want to zoom out before we go further, Ada, because there's a structural thing sitting under all of this we haven't said out loud yet: three of the four authors here are Lightmatter employees, and the paper explicitly promotes Lightmatter's own Passage platform. Every multiplier in that heatmap — including the eight-point-five-x headline number — comes out of a cost model these same authors built, control, and never validate against real photonic hardware, or even a real electrical cluster running the same workload. How much should that number actually be worth to us? 35 00:14:21,949 --> 00:14:59,194 [Dr. Ada Shannon] Honestly, less than the abstract wants it to be worth. This isn't independent replication — it's a vendor doing careful internal modeling of its own bet, useful information but not evidence a customer can spec hardware against. R4 makes that concrete: it's the single largest number in the whole paper, and it's still resting on that hypothetical chip we flagged earlier. The headline figure isn't measured against a real chip, or even a public roadmap chip — it's a projection stacked on another company's projection. Leading the abstract with that number, when real production platforms top out closer to four-and-a-half-x, is a choice, and not a neutral one. 36 00:14:59,194 --> 00:15:18,884 [Hal Turing] I'll push back on you there, though. Strip R4 out entirely and you're still left with real gains on shipping-adjacent hardware — two-point-one to five-point-eight-x depending on regime. That's not nothing, and it doesn't depend on a hypothetical GPU. Isn't focusing the critique on R4 letting the more defensible part of the claim off the hook a little too easily? 37 00:15:18,884 --> 00:15:58,776 [Dr. Ada Shannon] No, I actually disagree with you there, Hal — because R4 isn't the only inflation mechanism, it's just the most visible one. Every single cell in that heatmap, R4 or not, is the maximum multiplier across the whole device and batch sweep, not a typical operating point. The methodology section says it plainly: they take the max across all device counts, the max across all batch sizes. That's the best cherry from a wide tree, reported as if it were representative. An operator running at their actual, fixed batch size and device count would see something well below whatever number's on this table — for every platform, not just the speculative one. 38 00:15:58,776 --> 00:16:13,126 [Hal Turing] Wait, wait, hold on — that's actually a fair gut-punch, because it means even the 'safe' 2.1x number is already the best case, not the typical case. So then what should an infra architect actually walk away trusting here? 39 00:16:13,126 --> 00:17:08,343 [Dr. Ada Shannon] Treat the whole range as interconnect-sensitivity bounds, not a sizing rule. It tells you where communication is the bottleneck and how much headroom photonics could recover there — genuinely useful for prioritizing engineering attention. It does not tell you what your rack will do on Tuesday. Worth knowing too: this isn't Lightmatter's first paper making this argument. Bernadskiy, Carson, Graham, Groves and colleagues published 'Accelerating Frontier MoE Training with 3D Integrated Optics' in 2025, applying the same photonics thesis to training — this paper literally says it extends that analysis. Same vendor, overlapping authors, second installment. That's a research program, not outside confirmation. To their credit, the authors' own future-work section gets this right — they call for real optical hardware prototypes and full prefill-decode co-design, which is exactly the gap between what's claimed and what's shown. 40 00:17:08,343 --> 00:17:46,563 [Hal Turing] That's a good place to land it. So — real physics constraint, real simulated signal that photonics helps in communication-bound prefill, but the size of that signal is inflated by cherry-picked maximums and anchored, in the flashiest case, to hardware that doesn't exist yet, from a vendor grading its own product. Until someone runs this on actual optical silicon with decode sized to match, treat two-to-five-x as the honest range and the eight-x as a marketing asterisk. Thanks for breaking that down, Ada. And thanks to everyone listening — that's it for this one. See you next time.