DeepSeek-V4.1-Flash: Shrinking KV Cache at 552B Scale

A 552B-parameter MoE model whose KV cache footprint dropped ~437x since V1 — including a 4x jump-over-jump drop from V4-Flash — via a Causal Encoder-Decoder split, Sliding Window Attention, and CSA2 compressed sparse attention on FP4 storage.

arXiv:2609.19969 552B params · 8B/16B activated 890 bytes/token cache

Why the KV cache is the bottleneck

Long-context, tool-using agents are prompt-heavy: every processed token leaves behind a key/value pair per layer that must stay resident to avoid recomputing history. At million-token context, that cache — not compute — dominates serving cost.

Compression funnel (illustrative)

Four multiplicative levers stack to reach 890 bytes/token: entry-size compression (GQA/MLA), sequence compression (CSA/HCA), cross-layer sharing (CSA2), and FP4 storage.

Bars show relative bytes/token remaining after each lever is applied — magnitudes are illustrative of the mechanism stack, not disclosed per-stage figures.

437x since V1 (log scale)

DeepSeek reports ~437x cumulative reduction since V1, with ~4x of that arriving in this single generational jump (V4-Flash → V4.1-Flash).

Interim generation points (V2, V3) are log-interpolated for trend illustration — only V1, V4-Flash and V4.1-Flash values are paper-derived.

Pipeline: prompt to cached token

Causal Encoder-Decoder: halving prefill compute

40 layers split at the midpoint. The bottom 20 (encoder) compute their own KV from hidden state. The top 20 (decoder) project KV straight from the L/2 boundary hidden state via layer-dependent weights — no attention math required to build their cache.

Computes own KV (attention math) Projects KV from L/2 boundary (linear only)

Compute complexity

CSA2: three static modes per layer

Full Mode computes everything fresh. Reindex Mode reuses main KV but re-runs its own indexer query. Reuse Mode inherits both KV and the already-selected Top-K indices — no indexer compute at all. Hover a layer to see its assignment.

Full Mode Reindex Mode Reuse Mode

Hierarchical Sparse Indexer flow

A Full-Mode layer builds a shared candidate pool (~2,000 blocks × 8 positions ≈ 16k candidates) that later layers search or simply inherit — keeping later indexer cost constant instead of scaling with context length.

Table 1, re-plotted: base models, controlled eval

Params: total vs. activated

Claims vs. what's actually measured

Every headline byte-savings number derives from cache-entry arithmetic. No throughput curve, p99 latency table, or $/M-token figure appears anywhere in the paper.

MXFP4 storage format

E2M1 four-bit elements, one E4M3 scale shared per 16 channels — trained in via quantization-aware training, not patched on afterward. DeepSeek skips NVFP4's second global scale since post-RoPE channel magnitudes (≤~26) fit comfortably inside E2M1's range (up to ~2688).

SWA Bounded Replay: SSD vs. DRAM

Global KV keeps 72-hour SSD residency. Sliding-window KV moves entirely off SSD into a short-TTL pool using ~10% of host DRAM, accepting an approximate reconstruction on cache miss instead of exact replay.

Footprint vs. V4-Flash

References