← All episodes BASED: Balancing Recall and Throughput in Linear Attention

BASED: Balancing Recall and Throughput in Linear Attention

Sep 26, 2026
This episode dissects the BASED architecture from Arora, Eyuboglu, Zhang et al. (Stanford/Buffalo), examining how it tries to resolve the tradeoff between recall accuracy and inference throughput in sequence models. The discussion traces the lineage of alternatives to standard softmax attention — state space models like Mamba, gated-convolution approaches like H3 and Hyena, linear attention, and sliding window attention — explaining why a growing KV-cache makes attention memory-bound at scale, while fixed-size-state models trade away precise recall by construction. It highlights the MQAR benchmark from the Zoology paper as a tool for exposing this recall-versus-state-size Pareto frontier, showing that every architecture, including full attention, sits somewhere on that curve rather than escaping it. The hosts then unpack BASED's hybrid design, which combines a linear attention component for cheap global context with a small sliding window for exact local comparisons, and start evaluating whether this combination actually delivers on its claimed 24x throughput gain over FlashAttention-2. Listeners interested in efficient LLM inference will get a clear framework for why architecture choices around memory and recall are fundamentally linked rather than independently solvable engineering problems.
Sources:
1. Simple linear attention language models balance the recall-throughput tradeoff — Simran Arora, Sabri Eyuboglu, Michael Zhang, Aman Timalsina, Silas Alberti, Dylan Zinsley, James Zou, Atri Rudra, Christopher Ré, 2024
http://arxiv.org/abs/2402.18668
2. Zoology: Measuring and Improving Recall in Efficient Language Models — Simran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson, Michael Poli, James Zou, Atri Rudra, Christopher Ré, 2023
https://scholar.google.com/scholar?q=Zoology%3A+Measuring+and+Improving+Recall+in+Efficient+Language+Models
3. The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry — Michael Zhang, Kush Bhatia, Hermann Kumbong, Christopher Ré, 2024
https://scholar.google.com/scholar?q=The+Hedgehog+%26+the+Porcupine%3A+Expressive+Linear+Attentions+with+Softmax+Mimicry
4. Gated Linear Attention Transformers with Hardware-Efficient Training — Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, Yoon Kim, 2023
https://scholar.google.com/scholar?q=Gated+Linear+Attention+Transformers+with+Hardware-Efficient+Training
5. Mamba: Linear-Time Sequence Modeling with Selective State Spaces — Albert Gu, Tri Dao, 2023
https://scholar.google.com/scholar?q=Mamba%3A+Linear-Time+Sequence+Modeling+with+Selective+State+Spaces
6. Mistral 7B — Albert Q. Jiang et al., 2023
https://scholar.google.com/scholar?q=Mistral+7B