Speculative decoding speeds up dense-model inference by exploiting a memory-bandwidth bottleneck. Mixture-of-Experts routing breaks the assumption that every draft token costs the same to verify. This page visualizes why — and how EcoSpec's cost-aware selector responds.
Every autoregressive step streams the entire model's weights from HBM to compute one token, then discards the state. Compute units sit idle waiting on weight traffic — the classic roofline bottleneck. Speculative decoding amortizes that weight load across several tokens per pass.
In a dense model every candidate token costs the same to verify, so speculative decoding only has to optimize acceptance odds. In an MoE model, candidate tokens can route to different experts — verification cost becomes variable, and a confidence-only selector can't see it.
Draft tokens happen to route to experts already being loaded for other candidates — verification is nearly free.
High-confidence tokens fan out across disjoint experts, forcing extra multi-megabyte weight fetches just to check guesses that looked great to a confidence-only scorer.
Each cell is one expert; color intensity is how often it's touched during decoding. Concentrated ("peaky") routing gives a cost-aware selector real reuse headroom. Diffuse routing — DeepSeek-V3.1's 256-expert layer — leaves little to exploit no matter how good the cost signal is.
EcoSpec doesn't just rank by acceptance probability. It ranks by probability divided by the marginal cost of the experts a candidate would newly touch — so it can knowingly demote the single most probable token for a cheaper sibling.
score(token) = P(token) / marginal_cost(experts touched)
Table 4's ablation shows cost accounting isn't just a knob for trading acceptance away — a better cost estimate (dynamic, per-step marginal cost vs. a static global average) can raise acceptance length and improve speedup simultaneously.
EcoSpec trades a small amount of acceptance probability for reduced expert-loading traffic — by construction, since it can pick a cheaper, slightly-less-likely node.
The routing predictor costs about 4ms per step; verification-time savings from avoided expert fetches run in the hundreds of milliseconds — an order-of-magnitude win.
Both EAGLE-3 and EcoSpec are latency-sensitive, low-batch techniques. At batch size 8, both fall below plain autoregressive throughput — production MoE serving batches well past 4 specifically to amortize expert loads across concurrent requests.