This episode explores Shampoo, a 2018 optimization algorithm from Google Brain and Princeton that brings second-order curvature information to neural network training without the prohibitive cost of a full preconditioner. The discussion traces the lineage from Newton's method through AdaGrad's diagonal and full-matrix variants, explaining how Shampoo exploits the natural tensor structure of neural network weights — keeping separate small preconditioners per dimension and combining them via Kronecker products rather than materializing an impossibly large matrix. The hosts debate how much weight to put on the algorithm's convergence proofs, which rely on online convex optimization theory even though real network training is highly non-convex, concluding that the theory offers a sanity check rather than a guarantee and that empirical performance does the real work of justification. Along the way, they place Shampoo alongside K-FAC as one of the major structure-aware curvature approximations in the optimization literature. Listeners interested in the tradeoffs between cheap first-order methods and expensive second-order ones will find a clear walkthrough of how Shampoo threads that needle.
Sources:
1. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018
http://arxiv.org/abs/1802.095682. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization — John Duchi, Elad Hazan, Yoram Singer, 2011
https://scholar.google.com/scholar?q=Adaptive+Subgradient+Methods+for+Online+Learning+and+Stochastic+Optimization3. Scalable Second Order Optimization for Deep Learning (Distributed Shampoo) — Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, Yoram Singer, 2021
https://scholar.google.com/scholar?q=Scalable+Second+Order+Optimization+for+Deep+Learning+%28Distributed+Shampoo%294. SOAP: Improving and Stabilizing Shampoo using Adam — Nikhil Vyas, Depen Morwani, Rosie Zhao, Itai Shapira, David Brandfonbrener, Sham Kakade, 2024
https://scholar.google.com/scholar?q=SOAP%3A+Improving+and+Stabilizing+Shampoo+using+Adam5. Muon: An optimizer for hidden layers in neural networks — Keller Jordan et al., 2024
https://scholar.google.com/scholar?q=Muon%3A+An+optimizer+for+hidden+layers+in+neural+networks6. Adam: A Method for Stochastic Optimization — Diederik Kingma, Jimmy Ba, 2015
https://scholar.google.com/scholar?q=Adam%3A+A+Method+for+Stochastic+Optimization7. Decoupled Weight Decay Regularization — Ilya Loshchilov, Frank Hutter, 2017
https://scholar.google.com/scholar?q=Decoupled+Weight+Decay+Regularization8. Optimizing Neural Networks with Kronecker-factored Approximate Curvature — James Martens, Roger Grosse, 2015
https://scholar.google.com/scholar?q=Optimizing+Neural+Networks+with+Kronecker-factored+Approximate+Curvature9. A Stochastic Quasi-Newton Method for Large-Scale Optimization — Richard Byrd, Samantha Hansen, Jorge Nocedal, Yoran Singer, 2016
https://scholar.google.com/scholar?q=A+Stochastic+Quasi-Newton+Method+for+Large-Scale+Optimization10. Shampoo: Preconditioned Stochastic Tensor Optimization — Vineet Gupta, Tomer Koren, Yoram Singer, 2018
https://scholar.google.com/scholar?q=Shampoo%3A+Preconditioned+Stochastic+Tensor+Optimization11. Online Convex Programming and Generalized Infinitesimal Gradient Ascent — Martin Zinkevich, 2003
https://scholar.google.com/scholar?q=Online+Convex+Programming+and+Generalized+Infinitesimal+Gradient+Ascent12. Online Learning and Online Convex Optimization — Shai Shalev-Shwartz, 2012
https://scholar.google.com/scholar?q=Online+Learning+and+Online+Convex+Optimization13. Attention Is All You Need — A. Vaswani, N. Shazeer, N. Parmar, et al., 2017
https://scholar.google.com/scholar?q=Attention+Is+All+You+Need14. A Unified Approach to Adaptive Regularization in Online and Stochastic Optimization — V. Gupta, T. Koren, Y. Singer, 2017
https://scholar.google.com/scholar?q=A+Unified+Approach+to+Adaptive+Regularization+in+Online+and+Stochastic+Optimization15. Understanding Deep Learning Requires Rethinking Generalization — C. Zhang, S. Bengio, M. Hardt, B. Recht, O. Vinyals, 2017
https://scholar.google.com/scholar?q=Understanding+Deep+Learning+Requires+Rethinking+GeneralizationInteractive Visualization: Preconditioned Optimization Without the Full Matrix Cost