Sparse MoE and the collapse it hides
A dense Transformer fires every parameter on every token. Sparse MoE swaps the feedforward block for a bank of experts plus a router that activates only the top-k. The whole system leans on that router — and when it fails, experts drift toward identical functions.
Collapse at the frontier, not just in toy models
The paper's own motivating sweep checks ten current SOTA MoE models, four billion parameters up past 120 billion. Collapse shows up in essentially all of them — reasoning and non-reasoning alike.
Reading routing off the weights already there
An eigenvector of a weight matrix is a direction the matrix only scales, never rotates. Those directions already exist the moment training finishes — the authors' bet is that expert specialization is already baked into them.
What actually moved on real benchmarks
SSMoE-FULL keeps every expert and only changes routing. SSMoE-Pruned additionally drops 25% of spectrally-redundant experts for memory savings. Toggle between the two to see where the gains — and the credit — actually come from.
An average is not a deployment decision
The eight-benchmark average climbs even while specific tasks — GSM8K and ARC-Challenge especially — regress hard under the pruned config. That gap lives in an appendix, not the abstract.