This episode explores MiniMax Sparse Attention, a long-context transformer design that aims to preserve dense-model quality at million-token scale while sharply reducing the quadratic compute and memory costs of standard attention. It explains how the method combines Grouped Query Attention with blockwise sparse retrieval: a lightweight Index Branch scores past context in blocks, forces a recent local block to stay visible, selects top-k candidate regions, and then lets a Main Branch run exact softmax attention only inside those chosen blocks. The discussion places the paper alongside Longformer, BigBird, Routing Transformers, MInference, and Native Sparse Attention, arguing that its main contribution is a simpler, more GPU-friendly routing scheme that could make sparse attention practical at deployment time. Listeners would find it interesting because it focuses on the real technical tension behind ultra-long-context models: whether this kind of sparse routing can reliably recover rare distant evidence, or whether it mainly wins through recency bias and careful systems engineering.
Sources:
1. MiniMax Sparse Attention — Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, Pengyu Zhao, 2026
http://arxiv.org/abs/2606.133922. Longformer: The Long-Document Transformer — Iz Beltagy, Matthew E. Peters, Arman Cohan, 2020
https://scholar.google.com/scholar?q=Longformer%3A+The+Long-Document+Transformer3. Big Bird: Transformers for Longer Sequences — Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Amr Ahmed, 2020
https://scholar.google.com/scholar?q=Big+Bird%3A+Transformers+for+Longer+Sequences4. Efficient Content-Based Sparse Attention with Routing Transformers — Aurko Roy, Mohammad Saffar, Ashish Vaswani, David Grangier, 2020
https://scholar.google.com/scholar?q=Efficient+Content-Based+Sparse+Attention+with+Routing+Transformers5. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention — Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Wenfeng Liang, Wangding Zeng, 2025
https://scholar.google.com/scholar?q=Native+Sparse+Attention%3A+Hardware-Aligned+and+Natively+Trainable+Sparse+Attention6. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie et al., 2025
https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints7. Optimizing Mixture of Block Attention — Guangxuan Xiao et al., 2025
https://scholar.google.com/scholar?q=Optimizing+Mixture+of+Block+Attention8. MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention — Huiqiang Jiang et al., 2024
https://scholar.google.com/scholar?q=MInference+1.0%3A+Accelerating+Pre-filling+for+Long-Context+LLMs+via+Dynamic+Sparse+Attention9. DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads — Guangxuan Xiao et al., 2024
https://scholar.google.com/scholar?q=DuoAttention%3A+Efficient+Long-Context+LLM+Inference+with+Retrieval+and+Streaming+Heads10. FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision — Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, Tri Dao, 2024
https://scholar.google.com/scholar?q=FlashAttention-3%3A+Fast+and+Accurate+Attention+with+Asynchrony+and+Low-precision11. FIER: Fine-Grained and Efficient KV Cache Retrieval for Long-context LLM Inference — Dongwei Wang et al., 2025
https://arxiv.org/abs/2508.0825612. SCBench: A KV Cache-Centric Analysis of Long-Context Methods — Yucheng Li et al., 2024
https://arxiv.org/abs/2412.1031913. Streaming Video Question-Answering with In-context Video KV-Cache Retrieval — Shangzhe Di et al., 2025
https://arxiv.org/abs/2503.0054014. Native Hybrid Attention for Efficient Sequence Modeling — Jusen Du et al., 2025
https://arxiv.org/abs/2510.0701915. Rope to Nope and Back Again: A New Hybrid Attention Strategy — Bowen Yang et al., 2025
https://arxiv.org/abs/2501.1879516. AI Post Transformers: Optimizing Mixture of Block Attention Through Statistical Theory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-18-optimizing-mixture-of-block-attention-th-214f91.mp317. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp318. AI Post Transformers: Mooncake for KV Cache-Centric LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-05-mooncake-for-kv-cache-centric-llm-servin-1086d0.mp319. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp320. AI Post Transformers: δ-mem and Online Memory for LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-d-mem-and-online-memory-for-llms-6622fa.mp321. AI Post Transformers: Ministral 3: Cascade Distillation for Long-Context Multimodal Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-cascade-distillation-for-long-context-mu-0ebd1a.mp3Interactive Visualization: MiniMax Sparse Attention at Million-Token Scale