AI Post Transformers • Interactive Visual Companion arXiv: 2511.02237

Batch-Aware Expert Routing for Faster MoE Decoding

Visualizing why sparse MoE decode can still be slow: each token uses only a few experts, but a small batch can still activate a large union of experts, forcing irregular weight loads and making decode memory-bound.

Paper idea
OEA
No retraining
✓
Reported MoE decode gain
39%
Batch example
16

Core move: keep each token’s top-k₀ experts as a quality floor, then refill remaining slots using lower-ranked experts that are already active elsewhere in the batch.

baseline experts opportunistic piggyback newly loaded / expensive hot memory pressure

Why MoE decode becomes memory-bound

Dense FFNs reuse one shared weight block. MoE decode touches expert-specific weights that vary token-by-token. Small decode batches can still trigger a large union of experts.

token flow reused expert unique expert load memory bottleneck

Token-local sparsity ≠ batch-level sparsity

Hover cells: each row is one decode token, each column an expert. Even with top-2 routing, disagreement across tokens expands the batch-level union.

union = 11 experts

Two-phase routing rewrite

Step through the serving-time algorithm. The trained router ranking stays fixed; only the final selected set is edited to exploit already-loaded experts.

Per-token expert ranking matrix

Rows are tokens, columns are ranked experts. Toggle the preserved baseline k₀ to see how much quality floor is protected before opportunistic refill.

protected baseline per token = 2 experts

Reported component-level gains

Illustrative chart based on the episode discussion: OEA targets MoE-layer decode latency, not guaranteed end-to-end token latency.

Roofline-style intuition

Decode often sits in the memory-bound region. Reducing unique expert loads moves the operating point left-to-right more efficiently than token-local routing alone.

Batch-size sensitivity

OEA’s opportunity depends on who is in the batch. Small batches have little to piggyback on; moderate pooled decode tends to show the strongest effect.

Interactive batch composition simulator

Change batch size and overlap. The heatmap updates to show how many unique experts the batch activates under vanilla routing vs OEA.

B=16 • overlap=45%

References

Opportunistic Expert Activation: Batch-Aware Expert Routing for Faster Decode Without Retraining
Oncescu et al., 2025
arXiv:2511.02237
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Shazeer et al., 2017
Scholar link
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Fedus, Zoph, Shazeer, 2021
Scholar link
MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
Rajbhandari et al., 2023
Scholar link
PagedAttention
Kwon et al., 2023
Scholar link
The Roofline Model
Williams, Waterman, Patterson, 2009
Scholar link