This episode explores how speculative decoding can itself be pipelined, using the March 2026 paper "Speculative Speculative Decoding" to examine whether drafting and verification in LLM inference can be overlapped instead of run in a stop-and-go loop. It explains the bottleneck of standard autoregressive generation, reviews classic speculative decoding as introduced by Leviathan et al., and then focuses on the paper’s key idea: predicting likely verification outcomes so the next draft is ready before the verifier finishes. The discussion frames this as a scheduling and systems optimization problem rather than a new model architecture, connecting it to related work such as Lookahead Decoding, Medusa, and EAGLE. Listeners would find it interesting because it shows how careful inference-time execution design can deliver major practical speedups, including roughly a 30 percent average gain over strong speculative decoding baselines in the paper’s optimized Saguaro system.
Sources:
1. Speculative Speculative Decoding — Tanishq Kumar, Tri Dao, Avner May, 2026
http://arxiv.org/abs/2603.032512. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023
https://scholar.google.com/scholar?q=Fast+Inference+from+Transformers+via+Speculative+Decoding3. Break the Sequential Dependency of LLM Inference Using Lookahead Decoding — Yichao Fu, Peter Bailis, Ion Stoica, Hao Zhang, 2024
https://scholar.google.com/scholar?q=Break+the+Sequential+Dependency+of+LLM+Inference+Using+Lookahead+Decoding4. Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads — Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, Tri Dao, 2024
https://scholar.google.com/scholar?q=Medusa%3A+Simple+LLM+Inference+Acceleration+Framework+with+Multiple+Decoding+Heads5. Speculative Speculative Decoding — Tanishq Kumar, Tri Dao, Avner May, 2026
https://scholar.google.com/scholar?q=Speculative+Speculative+Decoding6. EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty — Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang, 2024
https://scholar.google.com/scholar?q=EAGLE%3A+Speculative+Sampling+Requires+Rethinking+Feature+Uncertainty7. Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding — Zachary Ankner, Rishab Parthasarathy, Aniruddha Nrusimha, Christopher Rinard, Jonathan Ragan-Kelley, William Brandon, 2024
https://scholar.google.com/scholar?q=Hydra%3A+Sequentially-Dependent+Draft+Heads+for+Medusa+Decoding8. Accelerating Production LLMs with Combined Token/Embedding Speculators — Davis Wertheimer, Joshua Rosenkranz, Thomas Parnell, Sahil Suneja, Pavithra Ranganathan, Raghu Ganti, Mudhakar Srivatsa, 2024
https://scholar.google.com/scholar?q=Accelerating+Production+LLMs+with+Combined+Token%2FEmbedding+Speculators9. Dynamic Speculation Lookahead Accelerates Speculative Decoding of Large Language Models — Jonathan Mamou, Oren Pereg, Daniel Korat, Moshe Berchansky, Nadav Timor, Moshe Wasserblat, Roy Schwartz, 2024
https://scholar.google.com/scholar?q=Dynamic+Speculation+Lookahead+Accelerates+Speculative+Decoding+of+Large+Language+Models10. Accelerating Large Language Model Decoding with Speculative Sampling — Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, John Jumper, 2023
https://scholar.google.com/scholar?q=Accelerating+Large+Language+Model+Decoding+with+Speculative+Sampling11. SpecInfer: Accelerating Large Language Model Serving with Tree-Based Speculative Inference and Verification — Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, Chunan Shi, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, Zhihao Jia, 2024
https://scholar.google.com/scholar?q=SpecInfer%3A+Accelerating+Large+Language+Model+Serving+with+Tree-Based+Speculative+Inference+and+Verification12. EAGLE-3: Scaling Up Inference Acceleration of Large Language Models via Training-Time Test — Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang, 2025
https://scholar.google.com/scholar?q=EAGLE-3%3A+Scaling+Up+Inference+Acceleration+of+Large+Language+Models+via+Training-Time+Test13. AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration — Bradley McDanel, 2025
https://scholar.google.com/scholar?q=AMUSD%3A+Asynchronous+Multi-Device+Speculative+Decoding+for+LLM+Acceleration14. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-Optimized+Large+Language+Model+Serving15. Layerskip: Enabling early exit inference and self-speculative decoding — not verified from snippet, 2024/2025
https://scholar.google.com/scholar?q=Layerskip%3A+Enabling+early+exit+inference+and+self-speculative+decoding16. Self-Speculative Decoding Accelerates Lossless Inference in Any-Order and Any-Subset Autoregressive Models — not verified from snippet, 2025
https://scholar.google.com/scholar?q=Self-Speculative+Decoding+Accelerates+Lossless+Inference+in+Any-Order+and+Any-Subset+Autoregressive+Models17. Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding — not verified from snippet, 2025
https://scholar.google.com/scholar?q=Speculative+Verification%3A+Exploiting+Information+Gain+to+Refine+Speculative+Decoding18. Parallelspec: Parallel drafter for efficient speculative decoding — not verified from snippet, 2025
https://scholar.google.com/scholar?q=Parallelspec%3A+Parallel+drafter+for+efficient+speculative+decoding19. Opt-tree: Speculative decoding with adaptive draft tree structure — not verified from snippet, 2024/2025
https://scholar.google.com/scholar?q=Opt-tree%3A+Speculative+decoding+with+adaptive+draft+tree+structure20. Cross-attention speculative decoding — not verified from snippet, 2025
https://scholar.google.com/scholar?q=Cross-attention+speculative+decoding21. Speculative Streaming: Efficient and Scalable Speculative Decoding with Multi-Stream Attention — not verified from snippet, 2025
https://scholar.google.com/scholar?q=Speculative+Streaming%3A+Efficient+and+Scalable+Speculative+Decoding+with+Multi-Stream+Attention22. Transactional KV Caching for Speculative Decoding under Paged KV Memory — not verified from snippet, 2025
https://scholar.google.com/scholar?q=Transactional+KV+Caching+for+Speculative+Decoding+under+Paged+KV+Memory23. Deft: Decoding with flash tree-attention for efficient tree-structured llm inference — not verified from snippet, 2025
https://scholar.google.com/scholar?q=Deft%3A+Decoding+with+flash+tree-attention+for+efficient+tree-structured+llm+inference24. AI Post Transformers: Adaptive Control for Batched Speculative Decoding in LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/adaptive-control-for-batched-speculative-decoding-in-llm-serving/25. AI Post Transformers: Accelerating Large Language Model Decoding with Speculative Sampling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/accelerating-large-language-model-decoding-with-speculative-sampling/26. AI Post Transformers: Apple's Speculative Streaming: Fast LLM Inference without Auxiliary Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/apples-speculative-streaming-fast-llm-inference-without-auxiliary-models/27. AI Post Transformers: QuantSpec: Hierarchical KV Cache for Self-Speculative Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/quantspec-hierarchical-kv-cache-for-self-speculative-decoding/28. AI Post Transformers: Memory Traffic Saturation in Transformer Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-20-memory-traffic-saturation-in-transformer-cd4961.mp3