AI Post Transformers — Episode Companion

Eigenvectors of Experts: Training-free MoE Routing Without Collapse

Giang Do, Hung Le, Truyen Tran · A2I2, Deakin University · ICML 2026
arXiv 2605.30992 Posted 2026-05-29 Sparse MoE routing Training-free Interpretability

Sparse Mixture-of-Experts models quietly waste capacity when different experts drift toward near-identical functions — representation collapse. This paper reads routing signal directly off the eigenvectors already latent in each expert's trained weight matrices, no retraining required. It works, with a proof behind it — but the validation doesn't fully cover the models that motivated the paper in the first place.

Sparse MoE and the collapse it hides

A dense Transformer fires every parameter on every token. Sparse MoE swaps the feedforward block for a bank of experts plus a router that activates only the top-k. The whole system leans on that router — and when it fails, experts drift toward identical functions.

Conditional computation: how a token moves through an SMoE block
Hover an expert to see its gate weight for this token. Only the top-2 experts (highlighted) are actually computed — the rest cost nothing for this token.
Expert output similarity — the fingerprint of collapse
Pairwise cosine similarity across 8 experts in one MoE layer. Under a purely learned router, experts converge toward each other (bright cells off-diagonal = collapse). Blending in the spectral signal pulls them apart.

Collapse at the frontier, not just in toy models

The paper's own motivating sweep checks ten current SOTA MoE models, four billion parameters up past 120 billion. Collapse shows up in essentially all of them — reasoning and non-reasoning alike.

Mean pairwise expert-output similarity across 10 frontier MoE models
Higher bars (redder) indicate more collapse. Hover a bar for the exact value. Illustrative values reconstructed from the paper's Figure 1 sweep description.
Gap flagged in Tab 5: Qwen3-30B, Qwen3-Next-80B, and ERNIE-4.5-21B appear in this sweep but are never tested with the paper's own fix.

Reading routing off the weights already there

An eigenvector of a weight matrix is a direction the matrix only scales, never rotates. Those directions already exist the moment training finishes — the authors' bet is that expert specialization is already baked into them.

From expert weights to routing logits — six stages
Click a numbered stage to see what happens there.
Lemma 3.1 — blending shrinks router-logit correlation
Drag the slider to set the balancing factor α. At α=0 this is vanilla SMoE; at α=1, pure spectral routing. Because eigenvector directions are close to orthogonal, blending them in provably lowers correlation between experts' logits.
α = 0.0 pairwise logit correlation ≈ 0.85
Sweet spot: empirically α ≈ 0.5–0.9 — the eigenvector signal carries most of the weight, but the learned router still gets a vote. Costs nothing extra: eigenvectors are computed once, offline.

What actually moved on real benchmarks

SSMoE-FULL keeps every expert and only changes routing. SSMoE-Pruned additionally drops 25% of spectrally-redundant experts for memory savings. Toggle between the two to see where the gains — and the credit — actually come from.

Reasoning benchmarks — % change vs. dense GPT-OSS-120B
SSMoE-FULL isolates the routing signal; SSMoE-Pruned bundles routing + a 25% expert cut. Bars below the line are regressions. Illustrative values consistent with the ~6% / ~6.4% averages and the reported GSM8K / ARC-C / OBQA / WinoGrande deltas.
8-benchmark average +6.4%
MTEB embedding quality — router baseline vs. MoEE vs. SSMoE
Same spectral analysis, repurposed as a free embedding extractor. SSMoE gains ~25–30% over a plain router baseline on the larger models; the gap narrows on the 1B model.

An average is not a deployment decision

The eight-benchmark average climbs even while specific tasks — GSM8K and ARC-Challenge especially — regress hard under the pruned config. That gap lives in an appendix, not the abstract.

The bundled headline vs. the per-task reality
GSM8K and ARC-Challenge deltas vs. dense GPT-OSS-120B, split by configuration. This is the regression the "+6%" headline is quietly averaging away.
Routing overlap: eigenvector vs. learned router
Share of tokens where the spectral router and the original learned router pick the same expert.
CLIP-MoE robustness under input corruption
Zero-shot retrieval/classification gain over baseline. SSMoE's edge grows once inputs are noisy — consistent with it being a genuinely different signal, not a rebrand of the same decision.
Verdict
Real yes: eigenvectors already latent in trained weights carry usable, provably decorrelating routing signal — Lemma 3.1 backs it mathematically, not just empirically.
Incomplete: the newest, most-collapsed models from the motivating sweep (Qwen3-30B, Qwen3-Next-80B, ERNIE-4.5-21B) are never validated; routing and pruning gains are bundled into one average; GSM8K is traded away for memory savings; MoEE is never run head-to-head on the reasoning suite. Worth watching, not worth deploying blind.

References

1Eigenvectors of Experts are Training-free Non-collapsing Routers — Giang Do, Hung Le, Truyen Tran, 2026 · arXiv:2605.30992
2Your mixture-of-experts LLM is secretly an embedding model for free — Li, Z. and Zhou, T., 2025 · Google Scholar
3On the representation collapse of sparse mixture of experts — Chi, Z. et al. (XMoE), 2022 · Google Scholar
4Dropping experts, recombining neurons: Retraining-free pruning for sparse mixture-of-experts LLMs — Zhou, Y. et al., 2025 · Google Scholar
5Small singular values matter: A random matrix analysis of transformer models — Staats, M., Thamm, M., and Rosenow, B., 2026 · Google Scholar
6SVD-LLM v2: Optimizing singular value truncation for large language model compression — Wang, X., Alam, S., Wan, Z., Shen, H., and Zhang, M., 2025 · Google Scholar
7Zero-shot sparse mixture of low-rank experts construction from pre-trained foundation models (SMILE) — Tang, A. et al., 2026 · Google Scholar
8Tight clusters make specialized experts — Nielsen, S., Teo, R., Abdullaev, L., and Nguyen, T. M., 2025 · Google Scholar
9Accuracy is not all you need — Dutta, A., Krishnan, S., Kwatra, N., and Ramjee, R., 2024 · Google Scholar