The paper’s core argument is not “bigger model, better model.” It is a systems claim: combine hybrid attention, sparse activation, and more stable residual transport so million-token context stops being a benchmark stunt and starts looking operationally plausible.
V4 shifts the burden away from “every token attends to everything at full cost.” The page below maps where long context becomes expensive, then shows how the paper splits the workload across selective retention, compressed global access, and sparse activation.
Illustrative runtime mix for a 1M-token request. Toggle to see the claimed move from flat full-history attention to partitioned long-context servicing.
The shape matters more than exact values. The point is how V4 tries to flatten the worst slope before 1M tokens.
CSA behaves like selective high-fidelity preservation. HCA behaves like lower-cost broad coverage. The visual goal is not exact reverse engineering, but to show why a hybrid can preserve needles without paying dense attention rent across the full million tokens.
Hover cells. Brighter cells mean more preserved attention mass across distance buckets and token classes.
Use the stage buttons to walk from raw context flood to query-conditioned retrieval over compressed memory.
Once MoE enters the picture, total parameters stop being the main serving metric. Active parameters per token, routing balance, and cross-layer signal transport start dominating the practical conversation. The mHC path then tries to keep depth usable rather than brittle.
Each column is a token family. Each row is an expert group. Hot cells indicate heavier routing demand.
mHC is shown here as a constrained multi-lane residual transport system, not just another skip arrow.
These charts emphasize the episode’s caution: the paper looks strong as an integrated system. It is weaker as a clean answer to which specific component earned each gain.
Older long-context families are plotted by qualitative retrieval strength and runtime tractability at very long windows.
Toggle between package efficiency and broader capability. The gap between those views is exactly the skepticism the episode argues for.
Compact links to the papers and prior episodes used to frame the systems story.