1 00:00:01,000 --> 00:00:44,746 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into "Simple Linear Attention Language Models Balance the Recall-Throughput Tradeoff" — the BASED paper. That's Simran Arora, Sabri Eyuboglu, Michael Zhang, and six co-authors, nine total, out of Stanford University and the University at Buffalo. It first hit arXiv in February 2024, revised in March 2025, and ran at the ICML 2024 workshop on efficient systems. And here's the number that got me: they claim 24x higher generation throughput than FlashAttention-2, on the same architecture family that's been trying to dethrone attention for years. 2 00:00:44,746 --> 00:01:10,845 [Dr. Ada Shannon] Right, and the reason that number matters isn't the number itself — it's what it's sitting on top of. This paper is basically an argument that recall and efficiency aren't two separate problems you solve independently. They're locked together. Every architecture that's tried to escape attention's cost has quietly paid for it somewhere, and this paper is the first one I've seen that draws the actual tradeoff curve instead of just claiming victory on one benchmark. 3 00:01:10,845 --> 00:01:21,759 [Hal Turing] Okay, let's back up for people who haven't been living in this literature. Why is attention even expensive in the first place? Because on paper, transformers are the thing everyone still reaches for. 4 00:01:21,759 --> 00:01:59,003 [Dr. Ada Shannon] So during generation, standard softmax attention has to keep a key and value vector for every single token you've generated so far — that's the KV-cache. It's exact, and it's why attention is so good at recall, meaning grounding what it says next in something specific you actually put in the context window earlier, not just something it memorized during training. But that cache grows linearly with sequence length. Generate a hundred thousand tokens and you're carrying a hundred thousand tokens' worth of memory around at every single step. That's the bottleneck — not the compute, the memory bandwidth of dragging that cache in and out. 5 00:01:59,003 --> 00:02:34,530 [Dr. Ada Shannon] So naturally people went looking for alternatives with a fixed-size state instead of a growing one. State space models — Mamba, from Albert Gu and Tri Dao out of Carnegie Mellon and Princeton in 2023 — compress everything into a small fixed recurrent state, like an RNN. H3, the Hungry Hungry Hippos paper from Fu et al. at Stanford in 2022, and Hyena from Poli et al. in 2023, do similar things with gated convolutions. RWKV does it too. They're fast, memory stays flat no matter how long the sequence gets, and that's fantastic for— 6 00:02:34,530 --> 00:02:53,152 [Hal Turing] Wait wait wait — hold on, Ada. If it's just a memory-versus-speed engineering tradeoff, isn't that basically a solved problem? Throw better hardware-aware kernels at it, IO-aware everything, and eventually the fixed-state models just catch up. That's what people said about a bunch of these Stanford Hazy Research papers before. 7 00:02:53,152 --> 00:03:21,202 [Dr. Ada Shannon] No — I actually disagree with you there, and this paper is specifically arguing against that read. It's not an implementation gap you engineer your way out of. If you squeeze a growing KV-cache into a small, fixed-size state, you are throwing away information by construction. It doesn't matter how clever your kernel is — a fixed-size state literally cannot hold an unbounded amount of precise, addressable information. That's not a software problem, that's a capacity problem. 8 00:03:21,202 --> 00:03:31,419 [Hal Turing] Okay but how do you even measure that in a way that isn't just vibes? Perplexity on some web corpus doesn't obviously tell you 'this model forgot the thing from ten thousand tokens ago.' 9 00:03:31,419 --> 00:04:17,301 [Dr. Ada Shannon] Exactly the gap they close. They lean on MQAR — multi-query associative recall — a synthetic benchmark that originated in the Zoology paper, Arora and Eyuboglu et al., Stanford, 2023, which is actually the same lab. You stuff a bunch of key-value pairs into the context, then later query specific keys and check whether the model retrieves the exact right value. It's clean, it's controllable, and critically, you can scale the number of pairs and the sequence length independently. That's how they build what they call the recall-memory Pareto frontier — plot recall accuracy against recurrent state size across architectures, and every single one of them, transformer included, lives somewhere on that curve. Nobody escapes it. You can only choose where on it you sit. 10 00:04:17,301 --> 00:04:24,732 [Hal Turing] So where does linear attention fit into that picture? Because that's the third option on the table here, right, next to full attention and SSMs. 11 00:04:24,732 --> 00:05:07,039 [Dr. Ada Shannon] Linear attention swaps out the exponential softmax kernel for a feature-map dot product — Katharopoulos, Vyas, Pappas, and Fleuret laid this out back in 2020, out of EPFL and Idiap, showing you could rewrite attention as a recurrence once you drop the exponential. That recurrence means fixed-size state, same as an SSM, but it's still fundamentally an attention mechanism doing a weighted average over the past rather than compressing it through gating. There's also sliding window attention, which just cheats differently — it's exact softmax attention, but only over the most recent w tokens, so your state is capped at the window width and anything older simply falls off a cliff. 12 00:05:07,039 --> 00:05:22,131 [Hal Turing] So global but coarse, versus local but exact. That's basically what BASED tries to have both of, isn't it — cheap linear attention keeping a rough thread on the whole sequence, plus a small sliding window doing precise nearby comparisons? 13 00:05:22,131 --> 00:05:30,816 [Dr. Ada Shannon] That's the pitch in a sentence. Whether that combination actually holds together under real workloads — and what it costs you — is exactly where we're headed next. 14 00:05:30,816 --> 00:05:40,754 [Hal Turing] Okay so walk me through the actual recipe, Ada. Not the pitch version — the mechanical one. What is literally happening inside a BASED layer when a token passes through it? 15 00:05:40,754 --> 00:06:22,828 [Dr. Ada Shannon] Three pieces, stacked. First, the global component: linear attention using that second-order Taylor approximation of softmax, running over the entire sequence so far — that's your cheap, unbounded-reach memory. Second, a tiny exact softmax attention window, only 64 to 128 tokens wide, sized specifically to keep tensor cores fully occupied rather than idle waiting on memory. Third, short gated convolutions — filter size 3 — threaded through the stack to handle fine-grained local token shifts. None of these three is new by itself. What's new is using all three together, each covering a gap the others leave open. 16 00:06:22,828 --> 00:06:32,116 [Hal Turing] Right, and that third piece is the one people skip over. Why do you need the convolution at all if you've already got exact local attention in the sliding window doing precise comparisons? 17 00:06:32,116 --> 00:07:15,445 [Dr. Ada Shannon] Because the sliding window only sees inside its own window — it doesn't touch the full sequence. The short convolution operates over the entire input, so it's doing a different kind of local shift, one that isn't bounded by window width. And this is exactly where the two-primitives-alone story breaks. Linear attention by itself fails at associative recall because that Taylor feature map, however clever, just doesn't have the precision for exact local token-by-token comparison — it's a smooth approximation, not a lookup. And sliding window attention alone caps your recall range at the window width — go wider and the state grows linearly with a genuinely nasty nonlinear hit to GEMM latency on tensor cores. 18 00:07:15,445 --> 00:07:28,634 [Hal Turing] Wait, hold on — that's three separate mechanisms doing three separate jobs. At what point does 'complementary building blocks' just become duct tape? Couldn't you glue together four more knobs and call it principled too? 19 00:07:28,634 --> 00:08:17,071 [Dr. Ada Shannon] No — I actually disagree with you there, Hal, because this isn't just empirical trial and error, it's backed by an actual theorem. They prove that any recurrent model depending causally on its input needs state size that scales with sequence length — order-N bits — to solve MQAR exactly. That's Theorem 3.1, and it's why the tradeoff isn't an implementation quirk you can engineer away, it's information-theoretic. Then separately they show BaseConv-style gated convolutions, the canonical stand-in for H3 and Hyena, provably need a growing number of layers, not a fixed constant-many, to solve the same task. That's why those architectures sit below the frontier in the plot we discussed. Each piece of BASED earns its place against a specific proven limitation. 20 00:08:17,071 --> 00:08:39,362 [Hal Turing] Okay, that's a fair correction — a theorem's a different bar than 'we tried it and it worked.' I'll grant the design isn't arbitrary. Though I'd still bet the hyperparameter search to tune three interacting knobs wasn't fun. Anyway — theory aside, does any of this survive contact with an actual GPU, or is linear attention back to being slow in practice? 21 00:08:39,362 --> 00:09:12,891 [Dr. Ada Shannon] That's the other half of the paper, and it's honestly the less flashy but more load-bearing part. Naive linear attention computes the feature map in Python and only offloads the dot product to CUDA — that's genuinely slower than a well-optimized softmax implementation, full stop. Their fix fuses the feature-map computation and the causal dot product into a single kernel, and critically, keeps the running KV-state in thread registers instead of shuttling it between HBM and SRAM on every tile. That's the difference between an elegant idea on paper and something you'd actually deploy. 22 00:09:12,891 --> 00:09:21,483 [Hal Turing] So does the elegant idea actually pay off? Give me the numbers against Transformer++ and Mamba — this is the part I actually care about. 23 00:09:21,483 --> 00:10:14,192 [Dr. Ada Shannon] At 1.3 billion parameters and 10 billion tokens: on SWDE, FDA, and SQuAD, Transformer++ scores 71.92, 73.23, 36.19. BASED gets 48.06, 24.41, 30.46. Mamba gets 34.74, 12.89, 28.20 — so BASED clearly beats Mamba on every recall task, though both trail full attention. Push to 50 billion tokens and the gap holds shape: Transformer++ 76.50, 80.47, 43.47; BASED 64.45, 30.40, 41.62; Mamba 52.75, 18.51, 35.92. That's the paper's headline — a 10.36 point average improvement over Mamba on downstream recall tasks at that scale, while matching it on overall perplexity and even nudging ahead on the LM-Eval common-sense average, 53.81 versus 53.50. 24 00:10:14,192 --> 00:10:17,629 [Hal Turing] And throughput — that's the number I keep seeing quoted everywhere. 25 00:10:17,629 --> 00:10:49,811 [Dr. Ada Shannon] Their IO-aware kernels give BASED 40 to 60 percent faster prefill than both FlashAttention-2 and Mamba at 4k sequence length. And on generation, the headline is up to 24 times higher throughput than FlashAttention-2 — that's specifically at 1.3 billion parameters, batch size 128, generating 1024 tokens, on a single H100. Concretely: 24.28 tokens per millisecond for BASED against 0.99 for Transformer++ under FlashAttention-2 at that setting. 26 00:10:49,811 --> 00:11:19,579 [Dr. Ada Shannon] But that number deserves scrutiny — it's one configuration, benchmarked against FlashAttention-2, a full-attention baseline with a KV-cache that grows with sequence length. Compare BASED against Mamba instead, the actual recurrent state-of-the-art it's supposed to be beating, and BASED only reaches about 95% of Mamba's throughput at that scale. At 360m parameters it's actually slower than Mamba at prefill in some configurations. The 24x is real, but it's answering the wrong comparison. 27 00:11:19,579 --> 00:11:41,128 [Hal Turing] That reframes the headline. It connects to something that bugged me in the recall numbers too — the paper frames this as "closing the gap" to attention, but the SWDE, FDA, and SQuAD numbers we covered earlier still show a 20 to 49 point gap to Transformer++, while the headline improvement over Mamba is only 10.36 points. 28 00:11:41,128 --> 00:12:04,487 [Dr. Ada Shannon] Fair tension — I'd rather be honest than wave it away. The gap does narrow with training, like we saw at 50 billion tokens, but even there FDA is still 50 points behind attention. "Closing the gap" is technically true relative to Mamba, but I'd push back on that phrase when the residual gap to real attention is still the dominant number on the page. 29 00:12:04,487 --> 00:12:35,323 [Hal Turing] Wait, hold on, hold on — that's exactly what I want to push on, because everything we just quoted is at 1.3 billion parameters and at most 50 billion training tokens. That's two or three orders of magnitude below where production language models live, and meaningfully below the token counts other 2024 seven-billion-parameter models were trained on. We know in-context learning shifts qualitatively with scale. So how much weight can we actually put on any of these curves holding up? 30 00:12:35,323 --> 00:13:08,899 [Dr. Ada Shannon] I actually disagree that it undermines the core claim — the empirical curves are scale-limited, sure, but Theorem 3.1 isn't. The Omega-N-bit lower bound on state size for exact recall is a communication-complexity argument; it doesn't care what parameter count you plug in. What's scale-limited is the empirical ranking — where BASED sits relative to Mamba at 70 billion parameters is genuinely unknown. But the tradeoff being real, not an implementation artifact you engineer away, that much the theory backs regardless of scale. 31 00:13:08,899 --> 00:13:37,181 [Hal Turing] That's a fair distinction, actually — the theorem grounds the tradeoff being real, it just doesn't tell us BASED's specific position on the frontier holds past 1.3B parameters. I'll take that correction. It's a real open question, not a solved one. Good, we agree on where the uncertainty actually lives. So let's talk about what's genuinely new here versus borrowed — where does the Taylor feature map actually come from? 32 00:13:37,181 --> 00:14:13,079 [Dr. Ada Shannon] It's not new — the Taylor feature map is from Zhang, Bhatia, Kumbong, and Ré's "Hedgehog and the Porcupine," Stanford, 2024. The recall-memory framework, MQAR, and the theory under Theorem 3.1 are lifted from Arora, Eyuboglu, Timalsina, Johnson, Poli, Zou, Rudra, and Ré's "Zoology," 2023 — James Zou is on both, and also behind "Recursive Multi-Agent Systems" from 2026. BASED's real contribution isn't a new primitive, it's the packaging — the window-plus-linear-attention combination and IO-aware kernels that make it fast. 33 00:14:13,079 --> 00:14:36,717 [Hal Turing] Which matters for how you read the comparison table, too — Gated Linear Attention, from Yang, Wang, Shen, Panda, and Kim, 2023, sits right there in Table 1 with dashes where the throughput numbers should be, because it wasn't benchmarked for efficiency. We don't actually know if BASED's edge is the Taylor map or just better kernel engineering than GLA had at the time. 34 00:14:36,717 --> 00:15:04,906 [Dr. Ada Shannon] Right. And there's a blind spot the paper never touches: attention's KV-cache isn't just a memory cost, it's the substrate an entire serving ecosystem is built on — PagedAttention-style block management, prefix caching for shared prompts, eviction schemes like StreamingLLM, post-hoc quantization. A compressed recurrent state can't be selectively evicted or shared across requests the same way. The paper never asks whether the efficiency win costs you all of that tooling. 35 00:15:04,906 --> 00:15:33,931 [Hal Turing] And the comparison set itself is already dated — Mamba, H3, Hyena, RWKV v4 and v5 reflect late 2023 into early 2024. Since then, Mamba-2, Gated DeltaNet, and RWKV-6 and 7 have gone after this exact recall weakness with more expressive state-update rules. "Expands the Pareto frontier beyond Mamba" is a snapshot against a baseline set that's already partly superseded, not a durable verdict. 36 00:15:33,931 --> 00:16:00,077 [Dr. Ada Shannon] So who should actually use this today? If you're serving at small-to-medium scale and memory-constrained — long context, resource-constrained inference — the tunable dial here, window size and feature dimension, is a genuinely useful knob. If you're running production-scale serving with a mature KV-cache stack already, the case is much weaker; you're trading a well-tooled ecosystem for savings that haven't been shown past 1.3 billion parameters. 37 00:16:00,077 --> 00:16:35,278 [Hal Turing] And the open question is obvious: does this hold at 7B-plus scale, and does the recipe generalize into newer gated-linear-attention hybrids. Wrapping up — BASED is a real demonstration that you can trade recurrent state size for recall on a dial, backed by an actual lower-bound theorem, not just vibes. But the framing oversells what's shown: the gap to real attention is still large, the throughput headline compares against the wrong baseline, and everything's validated well below production scale. Worth watching, not worth treating as settled. Thanks for listening.