Scout, Shortlist, Attend
MiniMax treats the KV cache like an internal retrieval corpus: block it, score it cheaply, keep the nearest block, then spend exact attention only on the shortlist.
Pattern Lab
Switch the matrix. Dense attention paints the whole causal triangle; MiniMax tries to preserve the useful parts while keeping the kernel shape regular enough to matter on real GPUs.
What the Speedup Means
The hard numbers in the episode sit on one side of the story. The other side is where the uncertainty still lives: cross-GPU portability and rare distant evidence recovery.
Where MiniMax Sits
Not the first sparse-attention idea, but one of the clearest attempts to strip routing down until deployment stops fighting back.