Shampoo already won AlgoPerf and shows up in how Gemini 1.5 Flash was trained — but it's expensive and finicky. SOAP proves Shampoo-with-half-power equals Adafactor inside Shampoo's own eigenbasis, then swaps in full Adam there. One new hyperparameter over AdamW. Claimed: 40%+ fewer iterations, 35%+ less wall-clock vs AdamW, ~20% better than Shampoo.
Each step trades exactness of the curvature estimate for tractability, until SOAP recovers full Adam-style adaptivity inside a rotated, curvature-aware basis.
Relative to a 100-unit AdamW baseline (lower is better). Toggle metric — the paper reports both.
For a weight matrix of shape m×n, illustrated here at m=8, n=6. Hover cells to see what they represent.
Click a stage or use the dots to step through the algorithm.
L, R (and their eigenvectors QL, QR) only get refreshed every f steps. Adam's second moment keeps updating every step regardless, even inside a stale rotation.
Final loss as eigendecomposition frequency f sweeps 1→100. Shampoo degrades fast; SOAP stays comparatively flat.
Steps-to-converge vs batch size. Dashed line = ideal linear scaling (2× batch → ½ steps).
210M / 360M / 660M non-embedding params, Chinchilla-optimal tokens. SOAP beats both baselines at every size shown.
SOAP is truncated at 0.5 / 0.625 / 0.75 / 0.875 of budget, then a 3-parameter curve a + bN−β is fit through those 4 losses and extrapolated to 1.0 to get the "40% fewer iterations" claim. No seeds, no error bars reported on this fit.
Mentions in SOAP's Jan 2025 revision, by competing/related method: