← All episodes TreeWY: Speculative Verification for Gated DeltaNet Hybrids

TreeWY: Speculative Verification for Gated DeltaNet Hybrids

Sep 3, 2026
This episode examines TreeWY, a proposed method for speculative decoding verification in hybrid language models that mix standard attention with Gated DeltaNet linear-attention layers. The discussion explains why current systems like vLLM and SGLang must snapshot the full recurrent state at every draft position before verification, since GDN's decay-and-overwrite state update can't be partially rolled back — a cost that multiplies across branches and makes wide speculative draft trees prohibitively memory-expensive. It traces the problem to its root, from the memory-bandwidth bottleneck that motivates speculative decoding in the first place to the mathematical mechanics of the gated delta rule that make hybrid-model states lossy and irreversible. The paper's proposed fix reframes the state update as decayed additive attention with a corrected value, hinting at a way to verify an entire draft tree with a single triangular solve rather than exhaustive snapshotting. Listeners interested in LLM inference efficiency will find the episode's central claim striking: a roughly 128x reduction in per-node memory without sacrificing correctness guarantees, potentially unlocking much more aggressive tree-based speculation on hybrid architectures.
Sources:
1. TreeWY: Speculative Verification for Gated DeltaNet Hybrids
https://arxiv.org/pdf/2608.20961
2. Bole: Efficient Tree Speculation for Hybrid-Attention Language Models — L. Wang et al., 2026
https://scholar.google.com/scholar?q=Bole%3A+Efficient+Tree+Speculation+for+Hybrid-Attention+Language+Models
3. ReplaySSM: Cache SSM Inputs, Not State — Dao AI Lab and NVIDIA, 2026
https://scholar.google.com/scholar?q=ReplaySSM%3A+Cache+SSM+Inputs%2C+Not+State
4. STree: Speculative Tree Decoding for Hybrid State-Space Models — Y. Wu et al., 2025
https://scholar.google.com/scholar?q=STree%3A+Speculative+Tree+Decoding+for+Hybrid+State-Space+Models
5. Parallelizing Linear Transformers with the Delta Rule over Sequence Length — S. Yang et al., 2024
https://scholar.google.com/scholar?q=Parallelizing+Linear+Transformers+with+the+Delta+Rule+over+Sequence+Length
6. Gated Delta Networks: Improving Mamba2 with Delta Rule — S. Yang, J. Kautz, A. Hatamizadeh, 2025
https://scholar.google.com/scholar?q=Gated+Delta+Networks%3A+Improving+Mamba2+with+Delta+Rule