Visual Podcast Companion

DeepSeek-V4 and Practical Million-Token Context

The paper’s core argument is not “bigger model, better model.” It is a systems claim: combine hybrid attention, sparse activation, and more stable residual transport so million-token context stops being a benchmark stunt and starts looking operationally plausible.

Interactive Viz URL Listen to the Episode Primary PDF 2026 • DeepSeek-AI • 1M-token context

Million-token practicality is a storage-and-routing problem

V4 shifts the burden away from “every token attends to everything at full cost.” The page below maps where long context becomes expensive, then shows how the paper splits the workload across selective retention, compressed global access, and sparse activation.

Inference Budget Flow

Illustrative runtime mix for a 1M-token request. Toggle to see the claimed move from flat full-history attention to partitioned long-context servicing.

active compute memory state compression / routing dominant cost

Cost Stack vs Context Length

The shape matters more than exact values. The point is how V4 tries to flatten the worst slope before 1M tokens.

Synthetic values anchored to the episode’s framing: V4 reduces relative KV and per-token inference burden at extreme context by changing what gets stored and what stays active.

Hybrid attention: keep a sharp local path and a cheaper global path

CSA behaves like selective high-fidelity preservation. HCA behaves like lower-cost broad coverage. The visual goal is not exact reverse engineering, but to show why a hybrid can preserve needles without paying dense attention rent across the full million tokens.

Attention Heatmap

Hover cells. Brighter cells mean more preserved attention mass across distance buckets and token classes.

cold / discarded medium / compressed hot / retained

Step-by-Step Context Pipeline

Use the stage buttons to walk from raw context flood to query-conditioned retrieval over compressed memory.

This diagram is intentionally mechanical: tokens, summaries, and recovered slots are separate objects because that is the real runtime story, not just a different equation on paper.

Sparse activation changes the relevant model-size number

Once MoE enters the picture, total parameters stop being the main serving metric. Active parameters per token, routing balance, and cross-layer signal transport start dominating the practical conversation. The mHC path then tries to keep depth usable rather than brittle.

Expert Routing Matrix

Each column is a token family. Each row is an expert group. Hot cells indicate heavier routing demand.

balanced share specialized route overloaded route

Residual Transport Diagram

mHC is shown here as a constrained multi-lane residual transport system, not just another skip arrow.

The visual claim: attention handles sequence interaction, while mHC tries to keep representations well-conditioned through depth so the long-context stack stays trainable and useful.

The strongest evidence is package performance, not clean causality

These charts emphasize the episode’s caution: the paper looks strong as an integrated system. It is weaker as a clean answer to which specific component earned each gain.

Long-Context Comparison

Older long-context families are plotted by qualitative retrieval strength and runtime tractability at very long windows.

The scatter is synthetic but disciplined: Transformer-XL, Longformer, and BigBird opened the design space; V4 is positioned as a more integrated million-token package.

Benchmark Shape

Toggle between package efficiency and broader capability. The gap between those views is exactly the skepticism the episode argues for.

Illustrative mock data. Bars are normalized to expose relative shape, not to reproduce a table from the paper.

References

Compact links to the papers and prior episodes used to frame the systems story.

DeepSeek-V4: PDF
DeepSeek-V3.2: Scholar search
Hyper-Connections: Scholar search
Muon is Scalable for LLM Training: Scholar search
LongBench v2: Scholar search
HyperRAG: Scholar search
ProphetKV: Scholar search
When Precision Meets Position: Scholar search
LongAttn: Scholar search
ResiDual: Scholar search
HyperAttention: Scholar search
AI Post Transformers: DeepSeek-V3: A Technical Report
AI Post Transformers: Ring-linear