A small "draft" model with a completely different tokenizer and architecture reads a 100,000-token prompt and tells a giant, unrelated target model which parts actually matter — cutting time-to-first-token by up to 18x, training-free.
Hover a box for the citation. Each step keeps the "small model proposes, big model consumes" shape but relaxes a constraint the previous step depended on.
Top strip: draft model's per-chunk importance score. Middle strip: target model's actual attention (ground truth) — correlated but noisier, illustrating why salience transfer is empirical, not guaranteed. Bottom strip: which chunks survive top-K selection + adjacent-chunk merging. Hover any cell.
DeepSeek ships no small sibling model — every cell here is a cross-family pairing by necessity, not by choice. Hover a cell for detail.
DS = DeepSeek (671B MoE). The LLaMA-1B → LLaMA-8B pairing is the same-family reference case Ada flagged: at only ~8x, the draft's own forward pass is a real slice of total cost — not the rounding error it is against a 671B MoE target.
128k-token prompt on SambaNova RDU. Compressing to 32k is the conservative setting (4.3s); 16k is where the headline 18x figure lives (2.5s).
* RULER's "full prompt" baseline already runs under SnapStream KV-cache compression on-chip — lossy vs. lossy, see Skepticism Scorecard. Code Debug at aggressive keep rate is the one clear counterexample to graceful degradation.
Hover a marker for the actual position taken in the episode.