Visualizing why sparse MoE decode can still be slow: each token uses only a few experts, but a small batch can still activate a large union of experts, forcing irregular weight loads and making decode memory-bound.
Core move: keep each token’s top-k₀ experts as a quality floor, then refill remaining slots using lower-ranked experts that are already active elsewhere in the batch.
Dense FFNs reuse one shared weight block. MoE decode touches expert-specific weights that vary token-by-token. Small decode batches can still trigger a large union of experts.
Hover cells: each row is one decode token, each column an expert. Even with top-2 routing, disagreement across tokens expands the batch-level union.
Step through the serving-time algorithm. The trained router ranking stays fixed; only the final selected set is edited to exploit already-loaded experts.
Rows are tokens, columns are ranked experts. Toggle the preserved baseline k₀ to see how much quality floor is protected before opportunistic refill.
Illustrative chart based on the episode discussion: OEA targets MoE-layer decode latency, not guaranteed end-to-end token latency.
Decode often sits in the memory-bound region. Reducing unique expert loads moves the operating point left-to-right more efficiently than token-local routing alone.
OEA’s opportunity depends on who is in the batch. Small batches have little to piggyback on; moderate pooled decode tends to show the strongest effect.
Change batch size and overlap. The heatmap updates to show how many unique experts the batch activates under vanilla routing vs OEA.