AI Post Transformers • Interactive Companion
arXiv: 2502.16982

Muon Is Scalable for LLM Training

A visual tour of the paper’s central tension: Muon looks like a major compute-efficiency win, but only as a hybrid recipe that combines orthogonalized momentum on matrix weights with AdamW on embeddings and 1D parameters, plus extra stabilization rules that appear to matter a lot at long training horizons.

Claimed FLOPs Ratio
0.519×
Scaling-law fit: Muon recipe reaches target loss at about 52% of AdamW training FLOPs.
Optimizer Shape
Hybrid
Matrix-shaped hidden-layer weights use Muon; embeddings, head, and 1D terms stay on AdamW.
Stress Point
Stability
Weight decay and per-parameter update scaling are presented as load-bearing for long-run bf16 training.
Large-Scale Anchor
Moonlight
3B active / 16B total MoE, trained on 5.7T tokens, used as a frontier-style evidence point.

Where Muon Actually Lives

The method is not a universal optimizer swap. This view maps which parameter families stay on AdamW versus where Muon’s orthogonalized momentum is applied, then overlays the systems path needed to keep that recipe viable at LLM scale.

Orthogonalized Updates on Matrix Weights

This section draws the optimizer’s shape bias. A raw momentum matrix is progressively transformed toward an orthogonalized update direction, while row/column energy and per-parameter scaling determine how aggressively each tensor moves.

Does the Frontier Shift?

A compute-efficiency claim is stronger than a prettier loss curve. The visuals below compare an AdamW baseline against a Muon-style recipe using mock scaling-law trajectories and a Pareto-style view that mirrors the paper’s framing.

Long-Run Stability Is Not Optional

The transcript’s sharpest technical point is that vanilla Muon may look fast early, yet magnitudes can drift upward over long horizons. This lab visualizes why added weight decay and update scaling can be the difference between a clean run and a slow-motion blow-up.

References