1 00:00:01,000 --> 00:00:52,455 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're looking at a paper called "Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training," from Yishun Lu et al. — four authors total, out of the Engineering Science Department at the University of Oxford, posted to arXiv in mid-May 2026. And Ada, here's the number that stopped me cold: on their test hardware, one of these fancy curvature-aware optimizers was taking 96 seconds per training step. AdamW, on the same hardware, was doing a step in about a second and a half. That's not a small tax, that's an optimizer that's basically unusable at that pace. 2 00:00:52,455 --> 00:01:28,492 [Dr. Ada Shannon] Right, and that gap is exactly the wall everyone doing second-order training runs into, Hal. What's genuinely interesting about this Oxford group is they don't touch the optimizer math at all. Every Kronecker factor, every update rule, stays mathematically identical to standard Shampoo or SOAP. What they rebuilt is the plumbing around it — where the state physically lives, when the expensive linear algebra actually runs, how nodes talk to each other. Their claim is direct: second-order optimization never caught on not because the algorithms were bad, but because nobody built the runtime to support them at scale. 3 00:01:28,492 --> 00:01:47,486 [Hal Turing] Okay, before we get into their system, Asteria, can you set the stage for people who haven't opened an optimizer config in a while? Every major model release I can think of — GPT, Llama, Mistral — they're all still running AdamW or something close to it. Why has first-order optimization won, this completely, for this long? 4 00:01:47,486 --> 00:02:34,344 [Dr. Ada Shannon] Because it's a trade that scales. AdamW only looks at the gradient — the first derivative of the loss — plus a running estimate of gradient variance per parameter. That's cheap: one or two extra numbers per weight, pure elementwise math, basically zero GPU overhead. The cost is statistical inefficiency — you need far more gradient steps, far more tokens, to reach a given loss, because you're ignoring how the loss surface curves. Second-order methods use that curvature directly, converging in fewer steps. The lineage runs from Martens and Grosse's K-FAC paper out of Toronto in 2015, to Gupta, Koren, and Singer's Shampoo paper from Google in 2018, to SOAP from Vyas, Morwani, Kakade and collaborators at Harvard in 2024. 5 00:02:34,344 --> 00:02:51,805 [Hal Turing] Oh wait, wait, wait — hold on. If you're tracking curvature instead of just the gradient, doesn't that mean the optimizer state itself has to be way bigger? Like, a full Hessian for a billion-parameter model would be a billion-by-billion matrix, right? That's not even remotely storable. 6 00:02:51,805 --> 00:03:41,914 [Dr. Ada Shannon] Exactly, and that's the whole ballgame. Nobody stores the true Hessian — it's infeasible, full stop. Shampoo and SOAP approximate it per weight tensor with two much smaller matrices, L and R in this paper's notation, one per tensor dimension — that's the Kronecker-factorization trick from K-FAC. It's still asymptotically worse than AdamW: AdamW's memory scales linearly with N, the parameter count, while Shampoo-style preconditioners scale closer to N-squared, and periodically need a matrix inverse root or eigendecomposition — an O(N-cubed) operation that doesn't parallelize like ordinary matrix multiplies and doesn't overlap cleanly with the rest of a training step. So real convergence gains per step collide with real infrastructure cost per step. That collision is the entire premise of this paper. 7 00:03:41,914 --> 00:04:04,252 [Hal Turing] Okay, but here's where I push back a little, Ada — isn't this basically just an infrastructure paper wearing a research-paper costume? They didn't invent a new optimizer. They moved memory around, added some asynchronous scheduling, relaxed how often nodes have to agree with each other. That feels like really elaborate systems engineering, not a new contribution to optimization. 8 00:04:04,252 --> 00:04:41,682 [Dr. Ada Shannon] I actually disagree with you there, Hal. That framing assumes the algorithm and the system it runs on are separable, and for methods like this they're not — an optimizer that only works with unlimited GPU memory and instant communication isn't practical, no matter how elegant its math is. Anil, Gupta, Koren, and Singer made basically this argument at Google in 2020 with Distributed Shampoo — their point was that Shampoo's biggest obstacle was systems, not math. This paper extends that diagnosis somewhere Distributed Shampoo doesn't reach: single-node, memory-constrained, heterogeneous setups, not just big homogeneous GPU clusters. 9 00:04:41,682 --> 00:04:55,614 [Hal Turing] Fair, fair — I'll grant that "systems problem, not math problem" is a real intellectual claim, not just an engineering afterthought. I guess what I want to see is whether their fix actually holds up under real numbers, not just under the framing. 10 00:04:55,614 --> 00:06:01,559 [Dr. Ada Shannon] That's the right instinct, and it's basically the shape of the rest of this paper. They call the specific failure modes "three physical walls" — a vertical capacity wall, where a single GPU just runs out of room without other GPUs to shard onto; an overlap disruption wall, where those cubic-cost matrix inversions and eigendecompositions stall the compute-communication overlap that frameworks like FSDP depend on, so the GPU just sits idle during the update; and a global consensus wall, where existing second-order methods insist on synchronous, full-state updates across every rank at every sync point, chaining fast NVLink traffic to slow inter-node InfiniBand traffic and killing any chance of topology-aware asynchrony. Put the three together — capacity, overlap, consensus — and you get why nobody's running Shampoo at LLM scale outside a paper. Their system, Asteria, is a runtime — not a new optimizer — built to route around all three. They test it two ways: single-node on a memory-constrained DGX Spark, and multi-node on GH200 clusters. 11 00:06:01,559 --> 00:06:09,314 [Hal Turing] Okay, so given those three walls, what's actually different about how Asteria attacks them? Not the philosophy — the mechanism. 12 00:06:09,314 --> 00:06:48,614 [Dr. Ada Shannon] It starts with where state physically lives. Asteria splits second-order state asymmetrically instead of treating it as one blob to offload. The Kronecker factor matrices — the ones getting updated every step from gradients — stay on the GPU, because that's where they're produced and consumed most cheaply. The inverse-factor matrices, the expensive derived quantities, live in CPU-accessible memory, either UVM-backed on something like DGX Spark or plain CPU DRAM. And there's an optional NVMe tier underneath that, moved through io_uring for asynchronous staging when even host memory gets tight. The logic is: match placement to the production-consumption pattern, not to a uniform 'offload everything' policy. 13 00:06:48,614 --> 00:07:08,239 [Hal Turing] Ada, be honest — doesn't that just sound like ZeRO-Offload? Ren and Rajbhandari's DeepSpeed paper out of Microsoft, 2021, already moved optimizer state to CPU DRAM to escape GPU capacity limits. This feels like the same trick pointed at a different tensor shape. 14 00:07:08,239 --> 00:07:37,171 [Dr. Ada Shannon] I actually disagree with you there, Hal. ZeRO-Offload works because Adam's state is flat, vector-valued, and updated with cheap element-wise math — you can shuttle it back and forth without thinking hard about it. Second-order state is structured matrices that need an O(d^3) inverse-root computation, and that computation has to happen somewhere specific, on a specific processor, at a specific point in the lifecycle. You can't just apply a generic offload policy to that and expect it to work. 15 00:07:37,171 --> 00:08:08,008 [Dr. Ada Shannon] And that's exactly what the next piece shows — the engineering, not just the placement decision. They call it the hook-orchestrated shadow-state pipeline. Asteria hooks into FSDP's existing forward and backward hooks — not to change the training graph, just to get a scheduling signal. When a hook fires, it dispatches the O(d^3) inverse-root computation to CPU worker threads and stages the result transfer on a low-priority shadow CUDA stream, completely separate from the primary stream running your actual forward and backward passes. 16 00:08:08,008 --> 00:08:14,463 [Hal Turing] Wait, hold on — so the shadow stream never competes with the main training stream for GPU cycles at all? 17 00:08:14,463 --> 00:08:45,113 [Dr. Ada Shannon] Exactly, that's the whole point — it's low-priority by design, so it only fills in slack the scheduler already has. And on the distributed side they pair that with bounded-staleness selective coherence. Instead of syncing every preconditioner block every step, Asteria tracks per-block freshness in a coherence registry, and only blocks that exceed a staleness budget S get synchronized — first within a node, then across node representatives, hierarchically. Everything else is treated as a cache hit and skipped entirely. 18 00:08:45,113 --> 00:08:49,571 [Hal Turing] Does any of that actually show up in the numbers, or is it still theoretical elegance at this point? 19 00:08:49,571 --> 00:09:31,832 [Dr. Ada Shannon] It shows up hard. On DGX Spark, native KL-Shampoo has step-time spikes hitting 96 seconds every ten steps from those synchronous eigendecompositions. Asteria-KL-Shampoo flattens that to 66.5 seconds, with total overhead over AdamW down to just 1.2 to 2.6 seconds per step. And the energy telemetry backs it up — native SOAP runs at 119.7% of AdamW's SoC energy, Asteria brings that to 113.8%; KL-Shampoo goes from 117.1% down to 107.9%. Asteria-KL-Shampoo also posts the best loss-reduction-per-energy score of any method tested, 3.3010. 20 00:09:31,832 --> 00:09:36,661 [Hal Turing] And that holds once you scale out to actual multi-node clusters, not just one Spark box? 21 00:09:36,661 --> 00:10:34,211 [Dr. Ada Shannon] On GH200 nodes, yes. Asteria tracks native SOAP and KL-Shampoo almost exactly in step-wise convergence — no quality loss — but reaches the same loss in less wall-clock time because the exposed overhead is gone. Sweeping the staleness budget S shows 1 is too tight, but by S equals 3 to 5 the runtime gain plateaus with no measurable hit to final eval loss — in fact final eval loss barely moves across S from one all the way to ten — so they settle on S equals 5 as the number they carry forward into the bigger 1B and 7B runs. That holds scaling up to 1B and 7B models, and their strong-scaling runs from 2 to 16 nodes show Asteria consistently beating native baselines on both speedup and per-step time, especially for KL-Shampoo. That's the clean part of the story: parity in convergence, real wall-clock wins on GH200. But now that we've been through the actual numbers, I want to go back and pull on a few threads in Section IV that bugged me on a second read. 22 00:10:34,211 --> 00:11:09,861 [Hal Turing] Let's do that, because two things jumped out at me too. First — they build this whole three-tier memory story: GPU, CPU, and an optional NVMe tier with io_uring staging, JIT paging, packed triangular SPD storage, presented like a headline contribution. But I went through every DGX Spark and GH200 result in Section IV, and I can't find one number produced with NVMe staging actually enabled. Second, those energy figures in Figures 6 and 7 — is that really a solid foundation for the loss-per-energy claims? 23 00:11:09,861 --> 00:11:49,661 [Dr. Ada Shannon] On the first one — you're right, and it's worth saying plainly: the NVMe path is architecturally described, fully implemented in the C++ extension, and never exercised in a single reported experiment. Every Spark number is UVM-CPU-DRAM, every GH200 number is host memory. It's a capability, not a demonstrated one. On the second, it's worse than it looks — that whole loss-reduction-per-energy story, Asteria-KL-Shampoo hitting 3.3010 against AdamW's 3.0942, comes from one run, one DGX Spark unit, fifty training steps. No stated seeds, no repetition, no variance reported anywhere. 24 00:11:49,661 --> 00:12:15,936 [Hal Turing] Oh wait, hold on — fifty steps? That's nothing against an actual pretraining run that goes for tens of thousands. So when I see SOAP's SoC energy drop from 119.7 percent to 113.8 under Asteria, or KL-Shampoo from 117.1 to 107.9, how do I know that's a real systems effect and not noise from wherever those fifty steps happened to land relative to warmup or thermal ramp? 25 00:12:15,936 --> 00:12:52,836 [Dr. Ada Shannon] You genuinely don't, not from what's reported. And it compounds with the learning-rate selection — they sweep ten-to-the-minus-four to ten-to-the-minus-two per optimizer and report only the run with the lowest validation loss, no mention of multiple seeds per configuration anywhere. So Asteria's wall-clock and loss advantage over native SOAP and KL-Shampoo could be genuine systems efficiency, or partly a favorable draw from best-of-sweep selection — the paper gives us no way to tell those apart. It's also why I'd have liked to see PyTorch Distributed Shampoo, Shi et al., largely Meta AI, 2023, actually benchmarked instead of just described as 'a representative systems baseline.' 26 00:12:52,836 --> 00:13:32,811 [Hal Turing] That's a real gap — the strongest existing systems competitor only gets discussed, never run head-to-head. Same with Deep Optimizer States, Maurya and colleagues out of Argonne National Laboratory, 2024 — that's the paper they credit for the async-CPU-worker idea itself, and they never quantitatively compare its interleaved offload against their own second-order-specific tiering. It's a natural ablation sitting right there, and they skip it. Which loops back to something bigger: they motivate this whole project with an RTX 3090, twenty-four gigs of VRAM, and then test exclusively on DGX Spark and GH200. 27 00:13:32,811 --> 00:13:52,061 [Dr. Ada Shannon] I actually disagree with you there, Hal. Motivating a problem with a 3090 isn't the same as claiming they tested one, and they say so themselves in the Discussion — narrow-PCIe discrete GPUs and tensor or pipeline parallelism 'remain to be explored.' That's more candor than most systems papers offer. 28 00:13:52,061 --> 00:14:15,861 [Hal Turing] Sure, but the abstract still leans on that commodity-GPU framing to sell urgency, and every result comes from platforms with unusually fast host-device bandwidth — exactly the condition that makes the shadow-pipeline overlap easy to hide. If PCIe is the real bottleneck for state movement, the spike-flattening they show might just not transfer to a 3090. 29 00:14:15,861 --> 00:14:51,936 [Dr. Ada Shannon] Okay, that part I'll concede — the overlap mechanism leans on bandwidth they haven't tested at the low end. I just don't think that makes it a bait-and-switch; owning the gap in the Discussion is different from overselling it. And it's not the only untested extrapolation — S equals 5 was swept on the 660M model in Figure 9 and then just reused as 'representative' at 1B and 7B without re-sweeping, and the async staleness schedule is wall-clock dependent, so two runs with identical seeds could genuinely diverge if a CPU job finishes milliseconds late. Nothing in the paper addresses that non-determinism. 30 00:14:51,936 --> 00:15:41,436 [Hal Turing] There's also a reuse blind spot worth flagging — those Kronecker factor matrices aren't just optimizer plumbing anymore in the broader literature; they're used as importance signals for pruning, Fisher-style regularizers for continual learning, model-merging pipelines. Asteria evicts inv_factor_matrices via madvise MADV_DONTNEED or pages them to NVMe the moment the current step consumes them. If you wanted that curvature state afterward, this design is actively hostile to you. And they never address checkpoint and restart semantics either — with CPU workers computing asynchronously and staleness bounds letting GPU-visible state lag, it's unclear a checkpoint captures anything resumable, which matters a lot at sixteen nodes where failures happen. 31 00:15:41,436 --> 00:16:19,136 [Dr. Ada Shannon] Both fair. So where does that leave the verdict? The runtime engineering — the asymmetric tiering, the shadow-stream decoupling, the topology-aware staleness — that's genuinely novel; nobody's applied ZeRO-Infinity-style tiering, Rajbhandari et al., Microsoft, 2021, specifically to matrix-structured second-order state before. What's incremental is the optimizer math itself — SOAP and KL-Shampoo, Lin et al., 2026, aren't new here. And it fits a pattern for this group — Lu and Armour's own prior work on Fisher-orthogonal projection, Oxford, 2026, shows they've been circling curvature-aware training for a while. Asteria reads like the systems companion to that agenda. 32 00:16:19,136 --> 00:17:00,961 [Hal Turing] Practically, if you're training on a single unified-memory box or a GH200-class cluster, this says you don't need to abandon second-order methods for capacity reasons anymore — it's a real option. On narrow-PCIe hardware or with activation recomputation changing how much GPU idle slack the shadow pipeline has to hide latency in, you're on your own; that's unexplored territory, not a solved one. The bigger takeaway is that whether second-order optimization becomes practical at LLM scale depends as much on runtime engineering — state placement, async compute, selective sync — as it does on the optimizer math itself. 33 00:17:00,961 --> 00:17:14,536 [Dr. Ada Shannon] That's the whole thesis in one line, and it's the right one. The math didn't need saving. It needed a runtime that stopped treating matrix-structured state like it was just another flat tensor to shove off to CPU. 34 00:17:14,536 --> 00:17:28,461 [Hal Turing] Great note to land on. That's "Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training" — thanks for breaking it down, Ada, and thanks for listening, everyone. We'll catch you next time.