AI Post Transformers // Episode Companion

SOAP: Stabilizing Shampoo's Second-Order Optimizer with Adam

⬡ arXiv:2409.11321 Vyas, Morwani, Zhao, Kwun, Shapira, Brandfonbrener, Janson, Kakade Harvard / Kempner Institute v2 · Jan 31 2025

Shampoo already won AlgoPerf and shows up in how Gemini 1.5 Flash was trained — but it's expensive and finicky. SOAP proves Shampoo-with-half-power equals Adafactor inside Shampoo's own eigenbasis, then swaps in full Adam there. One new hyperparameter over AdamW. Claimed: 40%+ fewer iterations, 35%+ less wall-clock vs AdamW, ~20% better than Shampoo.

Optimizer Lineage: Adagrad → Adam → Shampoo → SOAP

Each step trades exactness of the curvature estimate for tractability, until SOAP recovers full Adam-style adaptivity inside a rotated, curvature-aware basis.

Headline Efficiency Claim

Relative to a 100-unit AdamW baseline (lower is better). Toggle metric — the paper reports both.

Full budget numbers are extrapolated from a 4-point scaling-law fit — see the Scrutiny tab before taking these as measured facts.

What Each Optimizer Actually Stores

For a weight matrix of shape m×n, illustrated here at m=8, n=6. Hover cells to see what they represent.

One SOAP Step, Rotate → Adam → Rotate Back

Click a stage or use the dots to step through the algorithm.

Preconditioning Frequency f

L, R (and their eigenvectors QL, QR) only get refreshed every f steps. Adam's second moment keeps updating every step regardless, even inside a stale rotation.

Frequency Ablation

Final loss as eigendecomposition frequency f sweeps 1→100. Shampoo degrades fast; SOAP stays comparatively flat.

AdamW baseline Shampoo SOAP

Critical Batch Size Scaling

Steps-to-converge vs batch size. Dashed line = ideal linear scaling (2× batch → ½ steps).

AdamW SOAP Ideal (dashed)

Model Scales Tested

210M / 360M / 660M non-embedding params, Chinchilla-optimal tokens. SOAP beats both baselines at every size shown.

The 4-Point Extrapolation

SOAP is truncated at 0.5 / 0.625 / 0.75 / 0.875 of budget, then a 3-parameter curve a + bN−β is fit through those 4 losses and extrapolated to 1.0 to get the "40% fewer iterations" claim. No seeds, no error bars reported on this fit.

A 3-parameter fit through 4 points has almost no slack — the extrapolated gap is exactly as sensitive to noise as that implies.

Memory Footprint (per m×n layer)

Missing From The Conversation

Mentions in SOAP's Jan 2025 revision, by competing/related method:

Muon (Jordan et al., 2024) attacks the same problem with zero eigendecomposition and has real production adoption — it isn't cited once.

References