AI Post Transformers Visual Companion
Speculative Speculative Decoding
A visualization-first guide to how Saguaro overlaps drafting and verification, where the latency hiding helps, and where extra branch work starts turning into waste.
Pipeline Overlap
The key visual difference is not “more tokens become independent.” It is that the draft side stops idling while verification is in flight. The chart below compares classic speculative decoding with SSD on the same decode cycle.
Draft idle fraction
38%
Classic speculative decoding mock workload
Overlap recovered
24 ms
Draft-side work hidden under verification
Prepared branches
3
Bounded branch menu in this illustration
Hit rate
72%
Realized verification outcome already precomputed
Useful draft work
Verification window
Speculated follow-on branches
Wasted work
Saguaro Branch Engine
Step through one SSD cycle. Each state shows what the draft model prepares before the verifier returns. The branch that matches the actual acceptance length can hand off immediately.
Hover any branch node to inspect predicted probability, prep cost, and whether it survives the verifier.
Outcome Surface
Mock heatmaps show how verification-outcome predictability can change with draft quality, decode temperature, and branch depth. Hotter cells mean a higher chance the prepared branch is the one you actually need.
Columns represent predicted verification outcomes from early reject to full acceptance. Rows sweep workload conditions.
Fast Inference from Transformers via Speculative Decoding
Leviathan, Kalman, Matias, 2023
Scholar link