arXiv:2603.02631 ICLR 2026 SambaNova AI 8 authors

Cross-Family Speculative Prefill Cuts Long-Context Latency

A small "draft" model with a completely different tokenizer and architecture reads a 100,000-token prompt and tells a giant, unrelated target model which parts actually matter — cutting time-to-first-token by up to 18x, training-free.

Upasani, Raju, Li, Ji, Long, Wu, Thakker, Wang · Cross-Family Speculative Prefill: Training-Free Long-Context Compression with Small Draft Models
18x
TTFT reduction, 128k → 16k tokens (46s → 2.5s)
100k+
token prompts ranked by a model that never saw the target's tokenizer
90–100%
of full-prompt accuracy retained on most benchmarks after compression

Lineage: from decode-time speculation to cross-family prefill compression

Hover a box for the citation. Each step keeps the "small model proposes, big model consumes" shape but relaxes a constraint the previous step depended on.

Why "cross-family" is the whole point

Algorithm 1: prompt → compressed prompt

Token importance across a 100k-token document

Low importance
High importance

Top strip: draft model's per-chunk importance score. Middle strip: target model's actual attention (ground truth) — correlated but noisier, illustrating why salience transfer is empirical, not guaranteed. Bottom strip: which chunks survive top-K selection + adjacent-chunk merging. Hover any cell.

Draft → target accuracy retention (% of full-prompt baseline)

≤88%
98%+

DeepSeek ships no small sibling model — every cell here is a cross-family pairing by necessity, not by choice. Hover a cell for detail.

Draft overhead: how much does the "free lookahead" pass actually cost?

DS = DeepSeek (671B MoE). The LLaMA-1B → LLaMA-8B pairing is the same-family reference case Ada flagged: at only ~8x, the draft's own forward pass is a real slice of total cost — not the rounding error it is against a 671B MoE target.

Time-to-first-token vs. compression level

128k-token prompt on SambaNova RDU. Compressing to 32k is the conservative setting (4.3s); 16k is where the headline 18x figure lives (2.5s).

Accuracy retained by benchmark, conservative vs. aggressive keep rate

* RULER's "full prompt" baseline already runs under SnapStream KV-cache compression on-chip — lossy vs. lossy, see Skepticism Scorecard. Code Debug at aggressive keep rate is the one clear counterexample to graceful degradation.

Where Hal and Ada actually landed

Hal Ada Marker spacing = how far apart they ended up, not how far apart they started

Hover a marker for the actual position taken in the episode.

References