Interactive Episode Companion

The Sparsity Wall: What Reiner Pope Told Dwarkesh About MoE and Sparse Attention

This page treats sparsity as a map of cost pools. MoE sparsifies the FFN path, sparse attention sparsifies the KV fetch path, and the point of the episode is that those savings stack because they hit different bottlenecks.

Anchor Papers
Clark et al. 2022
Unified Scaling Laws for Routed Language Models
Attention Side
DeepSeek 2025
Sparse attention with sublinear decode traffic
Podcast Lens
Dwarkesh × Reiner Pope
April 29, 2026 blackboard framing
Transcript ID Scan
No extra explicit arXiv IDs
Pattern `DDDD.DDDDD` found only in provided source list, not the transcript body
Cost Pool A
FFN
Cost Pool B
QK / KV
Key Claim
Compound
Pedagogical Lens
Bandwidth

Two Independent Sparsity Levers

Dense transformers pay every token through the same full FFN and full-history attention path. The episode’s main visual argument is that MoE and sparse attention target different cost pools, so they can compound rather than substitute.

Hover any block or band for local metrics.

Serving Pipeline: Dense vs Split Sparsity

Active compute Dormant capacity Memory traffic

Per-Token Cost Pool Split

FFN-side Attention-side

Traffic Pressure by Context Length

Dense decode Sparse decode

Routed FFN: Quality from Dense Backbone × Expert Count

Clark et al. fit a two-axis picture. Increasing underlying dense size helps, increasing expert multiplicity helps, and the gains are real but not free: active compute, routing balance, and communication still matter.

Toggle model family, then inspect the heatmap and routing path.
Model Family
Interpretation
Routed models hold active compute flatter while total capacity grows through experts.

Loss Landscape Heatmap

Higher loss Better Best region

Router Dispatch Diagram

Selected Available

Total vs Active Parameters

Total Active per token

Sparse Attention: Moving Decode from Linear Pain toward Sublinear Fetch

Pope’s blackboard emphasis is memory traffic per generated token. Sparse attention matters when it reduces how much KV cache you need to touch as context length grows.

Switch between dense and structured sparse lookup patterns.
Pattern
Readout
Dense attention touches every previous token. Structured sparse attention keeps locality and a smaller global scaffold.

Attention Matrix Explorer

Cold Warm Hot

Bytes Fetched per Generated Token

Dense O(L) Sparse O(sqrt(L))

Decode Step Breakdown

KV reads Math

Compound Sparsity: Savings Stack Only If the Systems Story Stays Honest

The whiteboard version looks clean: route the FFN, sparsify the KV fetch path, keep quality, and lower serving pain. The practical caveat is that routing imbalance, batch shape, and locality can give some of the gain back.

Use workload toggles to see where the bottleneck moves.
Workload
Headline
Long-context decode is where sparse attention changes the economics most sharply.

2×2 Bottleneck Matrix

Healthy zone Wall

Cost Stack Before / After

Attention FFN Routing overhead

Step-by-Step Serving Path

Fast path Bandwidth pressure

References

Compact source set for the visuals and terminology.

Core Paper
Clark et al., 2022
Core Paper
DeepSeek, 2025
MoE Context
Scaling sparse FFN systems
Scaling Context
Dense compute-optimal baseline lens