AI Post Transformers · Episode Companion

SnapStream: Taming KV-Cache Eviction for Dataflow Accelerators

SnapStream fuses SnapKV's attention-based token selection with StreamingLLM's sink-plus-sliding-window into a single fixed-size cache, proven on a real 671B-parameter DeepSeek-R1 deployment running 128K-token context at over 1,800 tokens/sec on SambaNova's SN40L dataflow accelerators.

arXiv:2511.03092 SambaNova Systems 671B · DeepSeek-R1 · 16-way TP Full Interactive Viz ↗

Prefill / Decode Pipeline

Two phases, two bottlenecks. Hover the stages to see what limits each one.

Decode is memory-bound — little math relative to the bytes of KV cache moved every step, so cache size directly throttles time-per-output-token (TPOT).

KV Cache Growth, Unbounded

Cache size scales with sequence length — this is the pressure SnapStream relieves.

Two Ingredients, Fused

SnapStream isn't a new algorithm — it's SnapKV and StreamingLLM forced into one hardware-friendly cache.

Why This Wasn't Already Standard

Two deployment obstacles that accuracy benchmarks never had to face.

Continuous batching — requests in the same batch sit at different points in their own lifecycle, so there's no single clean moment to trigger compression across the batch.
Static-graph compilers — dataflow accelerators fix tensor shapes ahead of time; standard eviction slices, indexes, and re-concatenates buffers of varying shape every call.

Fixed-Size Cache Layout

Built once at prefill: sink tokens + SnapKV top-K middle + recent window. Hover a segment.

Illustrative proportions — segment sizes are configurable hyperparameters (L-sink, K, L-recent), not fixed ratios.

Ring Buffer vs. Naive Slicing

Decode only rotates the recent window. Step through writes — scatter into a fixed index, no gather, no concat.

Naive Slicing's SRAM Cost

32k cache, naive slice-and-concat vs SnapStream's ring buffer, on a 520MB SN40L socket.

323MB of 520MB (over 60%) of on-chip SRAM went to index bookkeeping alone, before a single attention FLOP — naive slicing needed 4.6× the memory of standard decode.

16-Socket Deployment: Prefill vs. Decode Parallelism

DeepSeek-R1's Multi-Head Latent Attention (MLA) changes how the compressed cache is sharded across the 16 SN40L sockets.

4.0×
lower KV memory
4.3×
higher throughput
2–5%
extra prefill latency

Throughput by Context Length (Table V)

Toggle context length — the ~4.2–4.3× gain holds steady across all three.

Max Batch Size Unlocked

Same memory budget, far more concurrent requests.

Consistency Across Context Lengths

Speedup ratio stays in a tight band — more convincing than any single headline number.

Infinity-Bench KV Retrieval: The Failure Mode

Needle-in-haystack retrieval is exactly where pooling-based top-K selection breaks.

87 → 59 exact match for Llama+SnapStream — a real third of retrieval accuracy lost. Still far ahead of StreamingLLM alone (2.4) and SAGE-KV (1.00), but not "minimal degradation" on this task.

LongBench-v2 on DeepSeek-R1

Beats naive truncation clearly — but still a real gap from full attention.

Hyperparameter Ablation (Table VII)

Sink tokens are load-bearing; top-K budget barely matters.

Reasoning Benchmarks Never Trigger Decode Eviction (Table III)

Generated tokens vs. DeepSeek's configured L-recent window (31,744). Neither benchmark fills it.

Decode-side eviction never engages on AIME24 or LiveCodeBench — those results stress-test prefill-side SnapKV almost exclusively, not the full pipeline under sustained generation.

Split Verdict

The systems result is completely honest. The accuracy claim is where "minimal degradation" gets stretched.

Systems: fully earned. 4× memory reduction, 4.3× throughput, on a real 671B production deployment. Nobody disputes this.
Accuracy: overstated in one spot. Minimal outside dense retrieval, real degradation inside it — the paper averages away its own failure mode.

References

1SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators — Jonathan Li et al., 2025arXiv:2511.03092
2Plasticine: A Reconfigurable Architecture For Parallel Patterns — Prabhakar et al., 2017Scholar ↗
3SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts — SambaNova, 2024Scholar ↗
4Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads — Abts et al. (Groq), 2020Scholar ↗
5In-Datacenter Performance Analysis of a Tensor Processing Unit — Jouppi et al. (Google), 2017Scholar ↗
6DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized LLM Serving — Zhong et al., 2024Scholar ↗
7Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference — Tang et al., 2024Scholar ↗
8InfLLM: Training-Free Long-Context Extrapolation with an Efficient Context Memory — Xiao et al., 2024Scholar ↗
9DeepSeek-V3.2-Exp: Boosting Long-Context Efficiency with DeepSeek Sparse Attention — DeepSeek-AI, 2025Scholar ↗
10Cartridges: Lightweight and General-Purpose Long Context Representations via Self-Study — Eyuboglu et al., 2025Scholar ↗
11LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction (SAGE-KV) — Wang et al., 2025Scholar ↗