1 00:00:01,000 --> 00:01:00,489 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Ada, today we're doing SOAP. Not the dish kind, though after reading this I kind of want a shower. This is "SOAP: Improving and Stabilizing Shampoo using Adam," by Nikhil Vyas, Depen Morwani, Rosie Zhao, Mujin Kwun, Itai Shapira, David Brandfonbrener, Lucas Janson, and Sham Kakade -- eight authors total -- out of Harvard University and the Kempner Institute at Harvard. The arXiv listing shows version two dated January 31st, 2025. And here's the number that stopped me cold: they claim over 40% fewer training iterations and over 35% less wall-clock time than AdamW, plus another 20% on top of that versus Shampoo itself. For an optimizer paper, that is not a rounding-error claim. 2 00:01:00,489 --> 00:01:42,982 [Dr. Ada Shannon] And that number is exactly why I want the appendix before the abstract, Hal. Because there's a real tension baked into this title. Shampoo already won -- the DistributedShampoo implementation took the AlgoPerf optimization benchmark outright, cutting wall-clock time 28% over the field, and a form of higher-order preconditioning shows up in how Google trained Gemini 1.5 Flash. So this isn't "here's a new optimizer nobody trusts yet." It's "the second-order method that already won is annoying to actually run, and we think we know exactly why, down to a proof." That's a much more interesting starting point than the usual optimizer paper with a new acronym and a cherry-picked loss curve. 3 00:01:42,982 --> 00:02:31,418 [Hal Turing] So let's set the stage on why anyone outside an optimizer research group should care. Training a modern language model isn't just about scaling parameters and tokens -- the optimizer decides how efficiently you convert compute into loss reduction. Shave 30 to 40% off wall-clock time for the same final performance, and at cluster-scale GPU prices that's real money and real calendar time on every run, not a one-time trick. That's the whole reason we've seen a wave of efficiency-focused optimizers the last couple years, and it's why a benchmark like AlgoPerf exists in the first place -- to make these efficiency claims fight each other head-to-head instead of just sitting in a table in someone's paper. 4 00:02:31,418 --> 00:03:18,230 [Dr. Ada Shannon] Right, and everything here traces back to one 2011 paper: Adagrad, by Duchi, Hazan, and Singer. Adagrad keeps a full preconditioner matrix H, built by summing the outer product of each gradient with itself, and it uses the inverse square root of that matrix to scale the update. Mathematically elegant, computationally insane -- for a weight matrix with millions of entries, H is a matrix with trillions of entries. Adam, from Kingma and Ba in 2015, is basically the practical fix: a diagonal approximation of Adagrad. Instead of the full matrix, it just tracks a running average of the gradient, and a running average of the squared gradient, per parameter. Cheap, fast, and it's why AdamW has been the default for basically a decade. 5 00:03:18,230 --> 00:03:27,889 [Hal Turing] Okay, so if Adam is the cheap diagonal shortcut, where does Shampoo fit in? Because I know it's second-order, but I don't have a clean mental picture of what it's actually storing. 6 00:03:27,889 --> 00:04:04,345 [Dr. Ada Shannon] So Shampoo, from Gupta, Koren, and Singer's 2018 paper, sits in between. Instead of one enormous preconditioner over the entire flattened weight vector, it exploits the fact that your weight is already a matrix, say m by n, and keeps two much smaller preconditioners: an m-by-m matrix L and an n-by-n matrix R. It approximates the full Adagrad preconditioner as a Kronecker product of those two, and updates the weights using the inverse fourth-root -- or in this paper's preferred version, inverse half-power -- of L and R. So instead of tracking one matrix the size of your entire-- 7 00:04:04,345 --> 00:04:16,326 [Hal Turing] Oh wait wait wait -- so instead of one gigantic matrix over the whole flattened weight vector, you're tracking two small matrices, one shaped by each dimension of the layer, and combining them? 8 00:04:16,326 --> 00:05:00,955 [Dr. Ada Shannon] Exactly, that's the whole Kronecker trick -- you get curvature information along both the input and output dimensions of the layer without ever materializing the full thing. The catch is that L and R still need periodic eigendecomposition to actually be used, which is expensive, and Shampoo picks up extra hyperparameters -- exponents, grafting schemes -- on top of that cost. This paper's core move is proving something specific: that Shampoo, run with the one-half power, is mathematically equivalent to running Adafactor -- Shazeer and Stern's 2018 memory-efficient Adam variant -- inside the eigenbasis that Shampoo's own preconditioner defines. Once you see it that way, the obvious next question is: why settle for Adafactor in that basis when you could run full Adam there instead? 9 00:05:00,955 --> 00:05:13,679 [Hal Turing] Which is literally the acronym -- ShampoO with Adam in the Preconditioner's eigenbasis, SOAP. So mechanically it's rotate into Shampoo's basis, but run Adam instead of Adafactor once you're there. 10 00:05:13,679 --> 00:05:49,903 [Dr. Ada Shannon] That's the shape of it, and it buys you something concrete: SOAP ends up with exactly one extra hyperparameter versus plain AdamW, called preconditioning frequency -- how often, in steps, you bother recomputing that expensive eigendecomposition. Everything else collapses back to familiar Adam machinery. And that's the headline claim we opened with: over 40% fewer iterations, over 35% less wall-clock time against AdamW, roughly 20% better than Shampoo on both. Worth remembering as we go: that's a claim we're going to want to poke at, not just repeat. 11 00:05:49,903 --> 00:06:07,828 [Hal Turing] Right, "machinery" is doing a lot of work in that sentence. So before we get to the payoff, give me the actual claim -- not the acronym, the theorem. What exactly are you proving is equal to what, and what does that buy you as an actual line of code in a training loop? 12 00:06:07,828 --> 00:06:58,262 [Dr. Ada Shannon] The claim is narrow: Shampoo, run with that one-half power we mentioned, produces the exact same update as rotating your gradient into the eigenbasis of Shampoo's own L and R matrices, running Adafactor's rank-one second-moment estimate in that rotated space, and rotating the result back. Same numbers, different description. Once you see Shampoo as a first-order method sitting in a second-order-derived basis, you're free to swap which first-order method goes inside. George et al. out of Montreal, NeurIPS 2018, did something structurally similar for KFAC with what they called E-KFAC -- a diagonal preconditioner updated between the expensive inversion steps. SOAP is that same move applied to Shampoo, except what drops into the rotated space is full Adam, not a plain diagonal average. 13 00:06:58,262 --> 00:07:21,436 [Hal Turing] Okay, so mechanically: every step you take the gradient, rotate it with Q_L transpose on one side and Q_R on the other, run Adam's first and second moment updates on that rotated gradient, rotate the result back with Q_L and Q_R, and take the step. And L and R themselves -- and their eigenvectors -- only get refreshed every f steps, that's the preconditioning frequency. 14 00:07:21,436 --> 00:08:03,975 [Dr. Ada Shannon] Exactly, and that's precisely the piece that fixes what kills Shampoo when you're stingy with compute. L and R accumulate every step, but the optimizer's actual adaptivity is frozen between eigendecompositions, since that's the only place curvature information gets baked in -- stretch that interval to save compute and Shampoo degrades. SOAP keeps updating its second-moment estimate, Adam's V, every single step even while sitting in a stale rotation, because a slightly outdated basis still beats no basis at all. They test this at 210 million, 360 million, and 660 million non-embedding parameters, Chinchilla-optimal token counts, and SOAP beats both AdamW and Shampoo at every size. 15 00:08:03,975 --> 00:08:22,922 [Hal Turing] So that forty-plus-percent-fewer-iterations, thirty-five-percent-wall-clock number you dropped earlier -- how do you actually compute that? You can't just eyeball two loss curves and declare a percentage, especially with a cosine decay schedule where the endpoint behavior is doing all the work and the two runs never finish at the same point. 16 00:08:22,922 --> 00:09:20,090 [Dr. Ada Shannon] Right, you can't just run SOAP for the same duration and read off a percentage, because the cosine schedule means final loss depends heavily on when you decay. So they run SOAP truncated at zero-point-five, zero-point-six-two-five, zero-point-seven-five, and zero-point-eight-seven-five of the training budget, take those four final losses, and fit a scaling law of the form a plus b times N to the minus beta through them -- that fitted curve is what gets compared against AdamW and Shampoo's full runs to extrapolate the crossover. Worth flagging as a modeling choice, not a stopwatch measurement. On the frequency ablation, they sweep from 1 to 100, and both optimizers beat AdamW throughout, but Shampoo's loss climbs much faster as frequency increases while SOAP stays comparatively flat. And on critical batch size, SOAP tracks the ideal linear-scaling line -- double the batch, halve the steps -- far more closely than AdamW does. 17 00:09:20,090 --> 00:09:43,031 [Hal Turing] None of this is free though, right? You're carrying L, R, their eigenvectors, plus the usual Adam moments, on top of the gradient and momentum buffers you'd already have with plain AdamW. What's the actual memory and compute bill here -- because if the answer is 'a lot more than AdamW,' that's a real adoption barrier at scale, not just an implementation detail. 18 00:09:43,031 --> 00:10:31,746 [Dr. Ada Shannon] For an m-by-n layer, SOAP needs two m-squared plus two n-squared plus three mn -- L, Q_L, R, Q_R, momentum, and the second-moment tensor, plus the gradient. That's identical to DistributedShampoo's footprint and noticeably more than AdamW's three mn. They propose closing that gap with a one-sided variant that projects only the smaller dimension and leaves the larger one as identity, and a factorized variant that swaps in Adafactor's rank-one estimate instead of full Adam inside the rotated space -- combine both and you actually undercut AdamW's memory. Compute-wise, the added per-step cost is m-cubed plus n-cubed plus two m-squared n plus two n-squared m over AdamW, measured via throughput on a single H100 using gradient accumulation to simulate larger batch sizes. 19 00:10:31,746 --> 00:10:49,765 [Hal Turing] Okay, here's what's bugging me about the headline numbers: four points, three free parameters in that scaling-law fit, and nowhere do I see multiple seeds or a confidence interval on the resulting 40% or 35% figure. That's one best-fit curve through four dots, extrapolated to 100%. 20 00:10:49,765 --> 00:11:18,093 [Dr. Ada Shannon] Right, and that's a real gap. A three-parameter fit through four points has almost no slack, so the extrapolated gap to AdamW's fixed endpoint is exactly as sensitive to noise as you'd expect. No error bars, no seed variation reported anywhere for these specific percentage claims. So when you read '40% fewer iterations,' what you're actually getting is one fit, one draw. Doesn't mean the number's wrong, but the precision implied by a clean percentage is more confidence than the methodology earns. 21 00:11:18,093 --> 00:11:47,536 [Hal Turing] That compounds with something they say themselves in the discussion, which I appreciate them being upfront about: the 210, 360, and 660 million parameter models are, quote, two orders of magnitude smaller than current LLMs, and generalizing to scale is explicitly called a hypothesis resting on theoretical foundation. Not evidence -- a hypothesis. Honest word choice, but it's doing a lot of work given how the abstract is framed. 22 00:11:47,536 --> 00:12:20,834 [Dr. Ada Shannon] And the kicker is they cite their own counterexample. Kaddour, Key, Nawrot, Minervini, and Kusner's 2023 paper, 'No Train No Gain,' out of UCL and Oxford, showed Lion, Sophia, and Adafactor-style optimizers looked great at small scale and then failed to beat a well-tuned AdamW at real LLM pretraining scale. SOAP's authors cite that paper as reason to distrust small-scale efficiency claims in general, while asking us to trust a similar unscaled claim for their own method. Not hypocrisy, exactly -- but the burden of proof is still unmet. 23 00:12:20,834 --> 00:12:27,335 [Hal Turing] Same pattern in the throughput numbers, right -- single-H100 proxy standing in for real large batches? 24 00:12:27,335 --> 00:12:59,100 [Dr. Ada Shannon] Exactly, and they say so directly -- that setup potentially disadvantages AdamW, since Shampoo and SOAP's per-step overhead gets compared against multiple accumulation steps instead of a real large batch. Their out is that overhead 'can be amortized across layers' in a distributed setting, the way DistributedShampoo does it. But that amortization is asserted by analogy, not measured for SOAP. The wall-clock number is a single-GPU proxy standing in for a distributed claim that was never actually run. 25 00:12:59,100 --> 00:13:15,215 [Hal Turing] Wait, hold on, that's the bigger issue for me -- if the whole efficiency story leans on 'trust us, it amortizes,' where's Muon in this conversation? That's a method attacking the exact same 'Adam is bad for matrix layers' problem, and it's already seeing real adoption. 26 00:13:15,215 --> 00:13:55,989 [Dr. Ada Shannon] Glaring omission, honestly. Muon, from Keller Jordan and collaborators, 2024, skips eigenbasis tracking entirely -- it orthogonalizes momentum with a few steps of Newton-Schulz iteration, cheap matrix multiplication, no eigendecomposition, no stored L, R, QL, QR at all. It picked up fast real-world traction -- open pretraining speedrun leaderboards, some larger runs -- specifically because the overhead is so much lower than Shampoo-family methods. SOAP's January 2025 revision doesn't mention it once. For a paper motivated by cutting LLM training cost, that's a notable blind spot. 27 00:13:55,989 --> 00:14:07,924 [Hal Turing] Same kind of gap in the memory story -- the combined variant is what closes it, not the base algorithm. But did they ever actually test low-precision storage for L, R, and Q? 28 00:14:07,924 --> 00:14:44,054 [Dr. Ada Shannon] No, that's the tell. They cite Wang, Li, Zhou, and Huang's 4-bit Shampoo paper, 2024, twice -- once for the power-iteration trick they borrow directly, once as future work for low-precision storage. SOAP itself never implements or tests quantized L, R, or Q. So 'less memory than AdamW' holds only for an untested combined variant, not the algorithm behind the headline numbers. Same story with Duvvuri, Devvrit, Anil, Hsieh, and Dhillon's 2024 Kronecker-approximation paper -- deferred to future work rather than compared against. 29 00:14:44,054 --> 00:14:46,609 [Hal Turing] So what should someone actually do with this today? 30 00:14:46,609 --> 00:15:14,705 [Dr. Ada Shannon] It's a low-risk swap. Nearly drop-in for AdamW, one new hyperparameter -- preconditioning frequency -- and at sub-billion-parameter scale, where they've actually tested it, the gains look real. If you're training small-to-mid transformers today, it's worth trying. The open questions the authors flag themselves are low-precision preconditioner storage, a real distributed implementation, and testing beyond language modeling into vision. None of those are solved yet. 31 00:15:14,705 --> 00:15:52,367 [Hal Turing] So here's where I land: genuinely clean theoretical contribution -- the Shampoo-equals-Adafactor-in-an-eigenbasis equivalence is real math, not marketing, and swapping in full Adam there is sensible. But the headline numbers are a four-point extrapolation with no error bars, the scale claim is an admitted hypothesis sitting next to the paper's own citation of a precedent where similar claims collapsed, and Muon -- cheaper, already adopted, solving the same problem -- goes completely unaddressed. Promising machinery, oversold headline. That's SOAP. Thanks for listening, everyone -- catch you next time.