AI Post Transformers Visual Companion

MiniMax Sparse Attention at Million-Token Scale

A million-token attention system drawn as routing geometry: one lightweight scout branch, a hard local keep, top-k remote block retrieval, and exact softmax only inside the chosen memory regions.

arXiv: 2606.13392 Posted June 11, 2026 109B multimodal MoE 3T training tokens 1M-token context target 28.4x less per-token attention compute 14.2x prefill, 7.6x decode on H800
Tab 1

Scout, Shortlist, Attend

MiniMax treats the KV cache like an internal retrieval corpus: block it, score it cheaply, keep the nearest block, then spend exact attention only on the shortlist.

Tab 2

Pattern Lab

Switch the matrix. Dense attention paints the whole causal triangle; MiniMax tries to preserve the useful parts while keeping the kernel shape regular enough to matter on real GPUs.

Tab 3

What the Speedup Means

The hard numbers in the episode sit on one side of the story. The other side is where the uncertainty still lives: cross-GPU portability and rare distant evidence recovery.

Tab 4

Where MiniMax Sits

Not the first sparse-attention idea, but one of the clearest attempts to strip routing down until deployment stops fighting back.

References

ArXiv Links and Related Threads

Comparison Set

Scholar links for the other papers discussed in the episode.

Related AI Post Transformers Episodes

Earlier public episodes used as comparison points, referenced by title.