1 00:00:01,000 --> 00:00:43,000 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper is "Eigenvectors of Experts are Training-free Non-collapsing Routers" — Giang Do et al., three authors total: Giang Do, Hung Le, and Truyen Tran, out of the Applied Artificial Intelligence Initiative, A2I2, at Deakin University in Victoria, Australia. It's an ICML 2026 paper, posted to arXiv on May 29th, 2026. Ada, the research question is blunt: can you route tokens to the right expert in a Mixture-of-Experts model with no learned router at all? 2 00:00:43,000 --> 00:01:04,650 [Dr. Ada Shannon] Normally 'training-free' on a systems paper makes my skepticism go up, not down — usually it means nobody had the budget to validate properly. What pulled me in here is the theory. They don't just show it works, they prove why it should reduce a specific failure mode mathematically, with a lemma behind it. That's rare — most routing papers just try five heuristics and report the winner. This one starts from an actual structural property already sitting inside the expert weights. 3 00:01:04,650 --> 00:01:26,075 [Hal Turing] Let's back up for anyone who isn't living and breathing Mixture-of-Experts architectures. What actually is Sparse Mixture of Experts, and why did the field pivot this hard toward it instead of just building bigger dense models? I know it's the default now for frontier models like GPT-OSS and DeepSeek, but walk me through what's structurally different. 4 00:01:26,075 --> 00:02:06,925 [Dr. Ada Shannon] A normal Transformer runs every token through the same dense feedforward block — every parameter fires every time. SMoE swaps that for a bank of smaller feedforward networks, the experts — could be 8, could be 128 — plus a small gating network, the router, that looks at each token and activates just the top-k, often top-1 or top-2. So a model might have 120 billion total parameters, but each token only touches a few billion of them. That's conditional computation — you scale capacity without scaling compute per token. It's how Mixtral, DeepSeek-MoE, and now GPT-OSS get away with being enormous while still shipping fast. The catch: the whole system leans entirely on that router making good decisions. 5 00:02:06,925 --> 00:02:25,250 [Hal Turing] And that's exactly where the paper's problem starts, right? A router is just a learned softmax over a projection — a pretty simple piece next to everything else in the stack. What happens when it doesn't learn to spread tokens out well? Because that's the whole game this paper's picking apart. 6 00:02:25,250 --> 00:03:01,125 [Dr. Ada Shannon] That's representation collapse — XMoE, Chi et al., 2022, laid it out theoretically first. Different experts, despite specializing separately, drift toward near-identical outputs. The router still fires, tokens still go somewhere, but you're paying for 128 experts and getting the capacity of a much smaller number. Everyone since has tried to patch it from the router side — HyperRouter, StableMoE, a few others — and they help some, but every fix requires retraining from scratch or at least fine-tuning the router. On a 120-billion-parameter model, that's not a weekend project. 7 00:03:01,125 --> 00:03:16,474 [Hal Turing] Oh wait, hold on — that's actually the part that got me. Because collapse isn't just 'a bit less efficient than advertised,' right? If experts are converging, you can't point to a token anymore and say 'that one went to the medical-reasoning expert.' 8 00:03:16,474 --> 00:03:52,425 [Dr. Ada Shannon] Exactly, and that's the deeper cost. Once experts collapse, whatever clean specialization story you wanted — this one handles legal reasoning, this one handles arithmetic — falls apart, because they're all converging toward the same function. That matters most in domains like medicine, law, and economics, where full explainability isn't a nice-to-have, it's a regulatory requirement. You can't deploy a black box that quietly stopped being modular in a courtroom or a hospital and just point at a benchmark score. So collapse is a capacity problem and an interpretability problem at once, and this paper goes after both. 9 00:03:52,425 --> 00:04:12,775 [Hal Turing] So here's the part that surprised me — you'd think with all the attention router design has gotten over the past few years, the big frontier MoE models would've mostly solved this by scale and better training alone. But the paper's own motivating sweep says otherwise, and that's basically why this paper exists. 10 00:04:12,775 --> 00:04:47,524 [Dr. Ada Shannon] Right, and I want to flag — this is their motivating observation, not yet the paper's own contribution being validated. They looked across ten current state-of-the-art MoE models, from a few billion parameters up past 120 billion — GPT-OSS-120B, the Qwen3-MoE family, ERNIE-4.5 — and found router collapse showing up in essentially all of them, to varying degrees, reasoning and non-reasoning models alike. So this isn't some toy setting where collapse only shows up in undertrained models. It's present at the frontier, in models people are shipping right now. That's the gap they're setting out to fill. 11 00:04:47,524 --> 00:05:05,625 [Hal Turing] Okay, so that's the setup — collapse is real, persistent even at scale, and fixing it the old way means retraining. Before we get into their fix, give us the plain-English version: what's an eigenvector of a weight matrix, actually? 'Eigenvector routing' sounds intimidating. 12 00:05:05,625 --> 00:05:44,650 [Dr. Ada Shannon] Think of a weight matrix as a machine that stretches and rotates whatever vector you feed it. Eigenvectors are the special directions where it doesn't rotate the input at all — it just scales it, longer or shorter, without changing where it points. Every matrix already has these built in; they're not learned separately, they're already latent the moment training finishes. The authors' bet: if an expert specialized during training on, say, arithmetic tokens, that specialization should already be baked into which directions its weight matrix responds to most strongly. So instead of learning a brand-new router to guess what each expert is good at, why not just read it off the expert's own geometry? 13 00:05:44,650 --> 00:06:24,125 [Dr. Ada Shannon] Here's the mechanics. Each expert has two feedforward matrices, W1 and W2. They form two Gram matrices — A equals W1 times its own transpose, B equals W2 transpose times W2 — which guarantees real, mutually orthogonal eigenvectors, no complex numbers involved. For each expert they keep only the eigenvectors from A and B most aligned with the router's own embedding, average those into one spectral embedding — basically a fingerprint of what that expert's geometry responds to — and stack the fingerprints across all experts into one matrix, W-EV. Multiply an input by that matrix and you get routing logits, no gradient required. 14 00:06:24,125 --> 00:06:43,850 [Hal Turing] So now there are two opinions on where a token should go — the learned router, and this new spectral one built purely from geometry. Do they replace the router outright, or run both together? Because for anything already deployed and working, swapping it out entirely feels like a bigger risk than the framing suggests. 15 00:06:43,850 --> 00:07:06,725 [Dr. Ada Shannon] They blend, not swap. A balancing factor, alpha, interpolates token-by-token between the two — zero is vanilla SMoE, one is pure spectral routing. The sweet spot empirically sits between roughly zero-point-five and zero-point-nine, so the eigenvector signal carries most of the weight, but the learned router still gets a vote. And it costs nothing extra: you compute the eigenvectors once, offline, and the whole thing drops into the forward pass — no fine-tuning, no backprop, no new data. 16 00:07:06,725 --> 00:07:25,550 [Hal Turing] Wait, hold on — training-free plus beats the thing that was actually trained usually sets off alarm bells for me, like nobody validated it properly. Is there real math behind why blending in these directions helps, or is this just something that happened to work out on their particular test set? 17 00:07:25,550 --> 00:08:01,450 [Dr. Ada Shannon] There is — Lemma 3.1, and it's the strongest part of the paper for me. It proves that mixing in eigenvector-derived logits shrinks the pairwise correlation between router logits across experts, given those directions are close to orthogonal. That's not hand-waved — independently drawn high-dimensional vectors are nearly orthogonal almost by construction, a basic property of high-dimensional spaces. So when the learned router's logits are already highly correlated, blending in something close to orthogonal mathematically guarantees the combined logits correlate less. That's a proof, not a pattern they happened to notice in one run. 18 00:08:01,450 --> 00:08:20,825 [Hal Turing] Okay, that's the routing story covered. You mentioned earlier that expert dropping rides on this same spectral analysis, purely for memory savings — and I want the headline numbers too, not just the theory. What did they actually see when they ran this on real reasoning benchmarks, at real model scale? 19 00:08:20,825 --> 00:09:10,525 [Dr. Ada Shannon] Same eigenvector analysis flags which experts are spectrally redundant — pointing in nearly the same direction as another expert already in the model — so they drop twenty-five percent of them, directly inspired by Zhou and colleagues' 2025 work on retraining-free expert pruning. Across GPT-OSS-20B and GPT-OSS-120B on eight reasoning benchmarks, that's roughly a six percent average gain over the original dense models, with twenty-three to twenty-five percent memory reduction. But the number that actually assigns credit correctly is SSMoE-FULL — same router, zero pruning, every expert kept — which still gains about six-point-four percent on average. The routing signal is doing essentially all the work; the memory savings from dropping experts are a bonus riding along, not the source of the accuracy gain. 20 00:09:10,525 --> 00:09:29,550 [Hal Turing] Good, that isolates the variable properly — credit where it's due, that's a clean ablation. But this framework started life as an embedding-extraction idea, so I'm curious how it holds up outside pure reasoning, on retrieval and classification where there's no single right answer to grade against. 21 00:09:29,550 --> 00:10:22,875 [Dr. Ada Shannon] On MTEB, tested across OLMoE-7B and its 1B sibling, Qwen1.5-MoE-7B, and DeepSeekMoE-16B, against a plain router baseline, vanilla SMoE, and MoEE — that's Li and Zhou's 2025 embedding-extraction paper, the closest prior work here — SSMoE gains roughly twenty-five to thirty percent on the larger models. They also push into vision-language with CLIP-MoE on zero-shot image-text retrieval and classification: smaller but statistically significant gains, around one percent on clean data, closer to three percent once you corrupt the input images with noise. That corrupted result backs the orthogonality claim too — the eigenvector router overlaps with the original router's expert picks only about fifteen percent of the time. It's not the same decision through a fancier lens — it's a genuinely different signal, one that degrades less than SMoE or MoEE when the input gets noisy. 22 00:10:22,875 --> 00:10:41,375 [Hal Turing] That fifteen percent overlap is a nice number — genuinely complementary, not a rebrand of the same decision dressed up differently. One result doesn't fit that clean story though, and I noticed it buried in the tables: what actually happens on GSM8K specifically? 23 00:10:41,375 --> 00:10:58,201 [Dr. Ada Shannon] Under the pruned version, GSM8K performance actually drops relative to the original dense model — math reasoning takes a real hit even while the eight-benchmark average looks great across the board. That's worth sitting with, and we'll come back to exactly what it means next. 24 00:10:58,201 --> 00:11:19,076 [Hal Turing] Okay, let's dig into that, because 'a hit on math reasoning' could mean a rounding error, or something structural about what pruning actually removes. Appendix D.7 has the real number, right? Because if that's where this lives, the six percent average everyone will quote in a slide deck is hiding something a lot less flattering. 25 00:11:19,076 --> 00:12:10,351 [Dr. Ada Shannon] Structural, and worse than it first looks. Under the pruned SSMoE variant — the one that delivers the twenty-three percent memory savings — GSM8K comes in twenty-two point nine percent worse, relatively, than the original dense GPT-OSS model, and twelve point six points worse than SSMoE-FULL, which keeps every expert and only changes routing. That gap is the whole story: the six percent headline bundles two separate interventions, a routing change and a twenty-five percent expert cut, into one number. The dropping idea itself comes from Zhou et al.'s 2025 paper, 'Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs' — their premise is that spectrally redundant experts are safe to remove. GSM8K says otherwise for multi-step arithmetic. Spectral similarity and functional necessity just aren't the same measurement, and this paper never reconciles that. 26 00:12:10,351 --> 00:12:37,726 [Hal Turing] That tension shows up again in Table 1. For GPT-OSS-120B at ten-shot, ARC-Challenge drops twenty-four and a half percent versus the original dense model, OpenBookQA is down three point six, WinoGrande down half a point — even while the eight-benchmark average climbs. If I'm deploying this for something narrow, like a commonsense-reasoning surface, the averaged 'SSMoE wins' framing could actively mislead me about my specific task. 27 00:12:37,726 --> 00:13:06,051 [Dr. Ada Shannon] Right, and that's the lesson buried in the appendix rather than the abstract. An average across eight benchmarks is a fine summary statistic, but it's not a deployment decision. What the paper should report — and doesn't — is routing-only gains and routing-plus-pruning gains as separate rows, so you can pick the configuration matching your actual workload instead of inheriting whatever tradeoff produced the best-looking average. Right now you'd have to dig through Table 1's fine print yourself to notice ARC-C regressed by nearly a quarter. 28 00:13:06,051 --> 00:13:58,751 [Hal Turing] Here's what bugs me structurally. Figure 1 — their own motivating evidence — sweeps ten current SOTA models from four billion parameters past a hundred and twenty billion: Qwen3-30B, Qwen3-Next-80B, ERNIE-4.5-21B, the whole current generation. That's the evidence collapse is real and widespread. But when they validate the fix, SSMoE only gets tested on GPT-OSS-20B and 120B, OLMoE, Qwen1.5-MoE, DeepSeekMoE, and CLIP-MoE. Qwen3-30B, Qwen3-Next-80B, ERNIE-4.5-21B — the newest, most collapse-prone models in their own chart — never get the treatment. Does this fix the problem they spent Figure 1 proving exists, or just the problem in the models they had on hand? 29 00:13:58,751 --> 00:14:36,651 [Dr. Ada Shannon] Oh — hold on, that's exactly what bugs me about the related-work story too. There's a closer comparison sitting right there that they dodge: MoEE, 'Your Mixture-of-Experts LLM is Secretly an Embedding Model for Free,' Li and Zhou, 2025 — literally the paper SSMoE builds its embedding-extraction framing on. It shows up in the MTEB tables, fine. It's completely absent from the GPT-OSS reasoning benchmarks, the exact tables where SSMoE's headline accuracy numbers live. If MoEE is your closest prior work and you're claiming reasoning gains, you run it on the reasoning suite. You don't get to compare against it only where you're strongest. 30 00:14:36,651 --> 00:14:58,151 [Hal Turing] So stepping back — genuinely novel, or incremental dressed up well? 'Read the specialization off the expert weights instead of learning a separate router' is a clean idea, and that fifteen percent overlap number from earlier already said it's not just a rebrand. But novelty and completeness aren't the same thing — what's your honest read? 31 00:14:58,151 --> 00:15:38,026 [Dr. Ada Shannon] Novel on the idea — nobody had shown expert-weight eigenvectors carry usable routing signal before, and Lemma 3.1 gives it real theoretical footing, not just a good result on one seed. Where it's incomplete is coverage: test the newest, biggest, most-collapsed models from Figure 1; report routing and pruning gains separately instead of one bundled average; find task-aware pruning that protects the experts GSM8K clearly needs, rather than pruning by spectral redundancy alone; and run MoEE head-to-head on reasoning. None of that undoes the core finding — eigenvectors already latent in trained weights carry real semantic signal. It just means the framing is ahead of what's been validated. 32 00:15:38,026 --> 00:16:16,201 [Hal Turing] So here's where I land. The research question gets a real yes — you can route without a learned router, and the geometry already sitting in expert weights is doing genuine work, not noise that happened to correlate. But 'non-collapsing' and 'training-free upgrade' are doing a lot of lifting for a result that's uneven per task, validated on only part of the lineup that motivated it, and quietly trading GSM8K performance for memory savings most people won't notice until their own eval suite does. Worth watching, not worth deploying blind. Thanks for listening, everyone — we'll catch you next time.