Speculative decoding was sold on "verification is nearly free." This episode traces the lineage from Leviathan's original trick through Medusa and EAGLE-3, then shows why sparse Mixture-of-Experts targets break that assumption — and how EVICT prices tree size against real expert-loading cost to claw the speedup back, losslessly.
Click a stage to see what it changed. Each step builds directly on the one before it.
Same tree, same verification pass — but a dense target loads one weight set no matter how many candidates it checks. An MoE target loads a different expert set per branch.
Hover a leaf. Every candidate token at every depth can route to a different top-k expert set — verifying the tree means loading the union of all of them.
Past ~30 nodes both curves climb steadily on Qwen3-30B-A3B.
Across three MoE models, checking guesses costs more than making them.
Four steps, no hand-tuned knob. Click each to expand.
Expected accepted length over verification cost, maximized to pick the tree size k.
Verification cost per tree size, measured once at model init — zero live timing at runtime.
Only Qwen3-30B-A3B reports a full autoregressive baseline; the other two are reported relative to EAGLE-3.
EVICT accepts fewer tokens per step — but cuts cost by far more.
ρ=1.0 is literally EAGLE-3 (keep everything). Every fixed threshold underperforms EVICT's adaptive utility.
The two closest relatives in EVICT's own related-work section never appear in the head-to-head comparison.
All headline numbers are single-stream. Under continuous batching, concurrent requests already touch most experts — the marginal cost of one tree's extra nodes shrinks toward zero.