A visual tour of the paper’s central tension: Muon looks like a major compute-efficiency win, but only as a hybrid recipe that combines orthogonalized momentum on matrix weights with AdamW on embeddings and 1D parameters, plus extra stabilization rules that appear to matter a lot at long training horizons.