SnapStream fuses SnapKV's attention-based token selection with StreamingLLM's sink-plus-sliding-window into a single fixed-size cache, proven on a real 671B-parameter DeepSeek-R1 deployment running 128K-token context at over 1,800 tokens/sec on SambaNova's SN40L dataflow accelerators.
Two phases, two bottlenecks. Hover the stages to see what limits each one.
Cache size scales with sequence length — this is the pressure SnapStream relieves.
SnapStream isn't a new algorithm — it's SnapKV and StreamingLLM forced into one hardware-friendly cache.
Two deployment obstacles that accuracy benchmarks never had to face.
Built once at prefill: sink tokens + SnapKV top-K middle + recent window. Hover a segment.
Illustrative proportions — segment sizes are configurable hyperparameters (L-sink, K, L-recent), not fixed ratios.
Decode only rotates the recent window. Step through writes — scatter into a fixed index, no gather, no concat.
32k cache, naive slice-and-concat vs SnapStream's ring buffer, on a 520MB SN40L socket.
DeepSeek-R1's Multi-Head Latent Attention (MLA) changes how the compressed cache is sharded across the 16 SN40L sockets.
Toggle context length — the ~4.2–4.3× gain holds steady across all three.
Same memory budget, far more concurrent requests.
Speedup ratio stays in a tight band — more convincing than any single headline number.
Needle-in-haystack retrieval is exactly where pooling-based top-K selection breaks.
Beats naive truncation clearly — but still a real gap from full attention.
Sink tokens are load-bearing; top-K budget barely matters.
Generated tokens vs. DeepSeek's configured L-recent window (31,744). Neither benchmark fills it.
The systems result is completely honest. The accuracy claim is where "minimal degradation" gets stretched.
| 1 | SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators — Jonathan Li et al., 2025 | arXiv:2511.03092 |
| 2 | Plasticine: A Reconfigurable Architecture For Parallel Patterns — Prabhakar et al., 2017 | Scholar ↗ |
| 3 | SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts — SambaNova, 2024 | Scholar ↗ |
| 4 | Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads — Abts et al. (Groq), 2020 | Scholar ↗ |
| 5 | In-Datacenter Performance Analysis of a Tensor Processing Unit — Jouppi et al. (Google), 2017 | Scholar ↗ |
| 6 | DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized LLM Serving — Zhong et al., 2024 | Scholar ↗ |
| 7 | Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference — Tang et al., 2024 | Scholar ↗ |
| 8 | InfLLM: Training-Free Long-Context Extrapolation with an Efficient Context Memory — Xiao et al., 2024 | Scholar ↗ |
| 9 | DeepSeek-V3.2-Exp: Boosting Long-Context Efficiency with DeepSeek Sparse Attention — DeepSeek-AI, 2025 | Scholar ↗ |
| 10 | Cartridges: Lightweight and General-Purpose Long Context Representations via Self-Study — Eyuboglu et al., 2025 | Scholar ↗ |
| 11 | LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction (SAGE-KV) — Wang et al., 2025 | Scholar ↗ |