1 00:00:01,000 --> 00:00:51,015 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper is Language Models Can Control Their Own Attention, from Namgyu Ho et al. — six co-authors in total — out of KAIST AI, with an advisory assist from Google DeepMind, posted to arXiv on September 2nd, 2026. Here's the number that grabbed me before I'd even finished the abstract: on a million-token conversation with a model like Qwen-3.5-397B, it has to haul roughly 15 gigabytes of key-value cache through memory for every single token it writes. Not once — every token. That's the same bandwidth bill as reloading the model's entire 17-billion active parameters, over and over. 2 00:00:51,015 --> 00:01:19,297 [Dr. Ada Shannon] And the maddening part is the model almost never needs all fifteen gigabytes of that read. Attention is famously peaky — a handful of tokens carry nearly all the weight, and the rest is dead freight you're still paying to move. What's interesting here isn't another scorer bolted on to guess which tokens matter. It's that they just ask the model. Have it say, in its own reasoning, 'here's where I need to look next.' Sounds almost too simple, which usually means it's either genuinely clever or quietly broken. 3 00:01:19,297 --> 00:01:43,539 [Hal Turing] That's the bet. Their framing question is basically: when you skim a long document to answer something, do you re-read every word each time? No — you remember roughly where the answer lives and go straight there. So why do transformers re-read the entire context from token one, at every decoding step, no matter what they're being asked? Said out loud, it's a strange way to have been doing this for years. 4 00:01:43,539 --> 00:02:20,133 [Dr. Ada Shannon] Worth defining for anyone who hasn't lived inside a transformer's internals: the KV cache is the running memory of every prior token's key and value vectors — it's what lets the model attend backward without recomputing from scratch. Generating each new token means pulling in the key and value for every token before it, weighting, summing. Trivial at a thousand tokens. At a million, that read stops being trivial and becomes the bottleneck. The matrix multiplies themselves are fast — it's dragging that much data off memory, every step, that eats your latency budget. 5 00:02:20,133 --> 00:02:58,632 [Hal Turing] So people already tried fixing this. The earliest approach was pure heuristic — keep what's recent, keep what's historically mattered, drop the rest. Guangxuan Xiao and Song Han's group at MIT, working with Beidi Chen and Mike Lewis at Meta AI, published StreamingLLM in 2023: they found the first few tokens of any sequence act as an 'attention sink' that soaks up outsized attention regardless of content, so you keep those plus a sliding window of recent tokens and the model streams indefinitely. Cheap, robust — and blind to what you're actually asking. 6 00:02:58,632 --> 00:03:34,577 [Dr. Ada Shannon] Then the field got smarter about query-dependence. Zhenyu Zhang's team, spanning UT Austin and CMU, published H2O in 2023 — it tracks cumulative attention scores per token and evicts everything except the heavy hitters and a recent window. Jiaming Tang and colleagues out of University of Washington and MIT then published Quest in 2024, scoring cache pages against your current query before deciding what to load. Both get filed under what this paper calls extrinsic scoring — a lightweight side mechanism, outside the model's own reasoning, guessing relevance for it. And the catch is— 7 00:03:34,577 --> 00:03:49,066 [Hal Turing] Oh — wait, sorry to cut you off, but that's exactly the quiet part: 'lightweight' still means touching every token or page's statistics before you can discard anything. That's still a full pass over the context, every single step. 8 00:03:49,066 --> 00:04:12,750 [Dr. Ada Shannon] Right — that's the wall every extrinsic method hits, dressed up differently each time. This paper's argument is almost insultingly simple: the model doing the reasoning already implicitly knows which part of the document is relevant, because that's just what competent reasoning over a document looks like. So instead of building a separate scorer to approximate that knowledge from outside, they just have the model say it out loud, in a format the inference engine can parse. 9 00:04:12,750 --> 00:04:51,202 [Hal Turing] Which pulls chain-of-thought back into this — the actual mechanism, not just 'let it think longer.' Jason Wei and colleagues at Google Research showed back in 2022 that prompting a model to externalize its intermediate reasoning as text before answering improves its answers, and as a side effect gives you a legible trace of what it's doing. This paper's trick, which they name Declarative Attention, repurposes that trace. Not just 'here's my reasoning toward the answer,' but 'here's my reasoning about which part of the context I even need to be looking at right now.' 10 00:04:51,202 --> 00:05:20,274 [Dr. Ada Shannon] I actually disagree that this is as obviously good an idea as it sounds, Hal. An extrinsic scorer, however crude, is at least an independent estimate — wrong sometimes, but not the same system grading its own homework. Here you're trusting the model's self-report of what it needs, and if it's overconfident or simply mistaken about that, the error doesn't show up as noise in some side channel. It becomes an actual hole in what the model can see when it writes the answer. 11 00:05:20,274 --> 00:05:36,760 [Hal Turing] Sure, but every method we just walked through is also a bet on an approximation being good enough — heuristics, proxy scores, none of it is ground truth either. Why is the model estimating its own relevance categorically worse than a side-model estimating relevance for it? 12 00:05:36,760 --> 00:05:54,779 [Dr. Ada Shannon] Because the failure modes differ, not because self-report is automatically worse — fair, I'll give you that much. But we don't have to settle it from the couch. The paper actually reports the accuracy hit this costs, and how it shrinks with model scale, and that's exactly where the numbers should do the arguing instead of us. 13 00:05:54,779 --> 00:06:20,924 [Hal Turing] Fair enough — that's exactly where we're headed. For now, here's the one-liner to hold onto: sparse attention has always meant guessing which context matters without paying to check it every step, and every method so far has done that guessing from outside the model. This paper's whole move is making it intrinsic — stop approximating from the outside, and just let the model tell you where it needs to look. 14 00:06:20,924 --> 00:07:03,835 [Hal Turing] making that guess something the model just says out loud in its own reasoning, instead of something an outside scorer infers for it. Take the example they actually use: the question is how long after founding did some company go public. The model opens in what they call global mode, skimming the whole context to decide where the founding date probably lives. Once it thinks it's found the spot, it drops into focus mode and reads only that one slice -- pulls out 2003. Goes back to global hunting the IPO year, focuses on a different chunk, gets 2011, and only then does the subtraction in local mode, attending to nothing but its own scratch work. 15 00:07:03,835 --> 00:07:52,690 [Dr. Ada Shannon] Local's the elegant part -- once both numbers are down, it doesn't need the source document anymore, just its own arithmetic, so why keep fifteen gigabytes of KV cache in view to compute 2011 minus 2003. But underneath that is a real engineering problem: how does a model "name" a chunk without you handing it arbitrary boundaries it's never seen before? Their fix is to split the context into roughly two-thousand-token "magic chunks" and present it as a fake tool-call transcript -- an assistant turn calls get_magic_chunk, a tool response comes back headed Magic Chunk 7, over and over, before generation even starts. No tool actually runs, it's theater, but it puts every segment boundary exactly on the turn-delimiter tokens the model already tracks fluently from post-training. 16 00:07:52,690 --> 00:08:10,662 [Hal Turing] Okay, but somebody still has to turn "focus on chunk seven" into a real memory operation the GPU respects, and vLLM stores KV cache in fixed blocks, not individual tokens you can cherry-pick. So how do you get from a text tag to an actual mask without rewriting the attention kernel? 17 00:08:10,662 --> 00:08:41,312 [Dr. Ada Shannon] You don't rewrite it at all, that's the trick. A small state machine sits alongside the inference engine, watches for the closing bracket on a focus or local tag, and flips the mask right then. Since vLLM only saves time by skipping whole blocks, it rounds each kept span outward to block boundaries -- a few dozen extra tokens at the edges, never dropping anything the model asked for. It's just a hook on vLLM's attention metadata builder rewriting the block table each step. FlashAttention runs completely unmodified; it just gets handed fewer blocks to read. 18 00:08:41,312 --> 00:08:47,721 [Hal Turing] So how do they actually test whether this holds up? "The model declares its own attention" is a nice story, I want numbers. 19 00:08:47,721 --> 00:09:38,015 [Dr. Ada Shannon] Six models across two families -- Gemma-4 at three sizes, Qwen-3.5 and 3.6 -- run against fifteen long-context sources from RULER, LongBench v1 and v2, LooGLE, and ZeroSCROLLS, from short passkey retrieval up past million-token code repos. Three arms: vanilla, DA, and a DA-no-mask ablation. Since free-form answers can't be exact-matched, they score with an LLM judge, Qwen-3.5-4B, checked against Gemini-3.1-Pro at Pearson r of 0.99. Headline numbers on the flagships: 52.0 percent fewer attended tokens on Gemma-4-31B, 31.1 percent on Qwen-3.6-27B, for accuracy drops of only 1.27 and 2.75 points. 20 00:09:38,015 --> 00:09:51,019 [Hal Turing] Oh -- wait, that no-mask arm, that's the control that isolates whether the savings come from the mask or just from writing in this chunked format, right? Because if the format alone did the work, the masking would be a much smaller deal than they're selling. 21 00:09:51,019 --> 00:10:34,208 [Dr. Ada Shannon] Exactly the control you'd want, and it lands cleanly in the mask's favor. DA-no-mask matches vanilla accuracy, but its attended tokens go up, 66 percent higher than vanilla on Gemma, because the DA-style prompting itself makes both arms run 15 to 35 percent more decode steps. The mask is what converts that overhead into a net win. And the scaling story matters more for where this goes: the accuracy gap to vanilla shrinks as the backbone grows, absolute token savings grow with context length too, up to 21 million tokens saved per response at the longest contexts. Translate that into roofline wall-clock and DA lands at 0.71x of vanilla decode time on Gemma, 0.77x on Qwen. 22 00:10:34,208 --> 00:10:44,610 [Hal Turing] A 30-ish percent wall-clock cut, on off-the-shelf models, zero training -- that's not a footnote, Ada, that's a genuinely deployable speedup if it holds. 23 00:10:44,610 --> 00:11:07,969 [Dr. Ada Shannon] I'd pump the brakes on "deployable." That 0.71 and 0.77 is a roofline number, charged at each operation's own hardware ceiling, 40 percent MFU, 70 percent MBU, on a single B200, and it excludes prefill entirely. It's a theoretical floor on what an optimized stack could achieve, not a stopwatch measurement under real batching pressure. 24 00:11:07,969 --> 00:11:34,951 [Hal Turing] Fair, it's a ceiling estimate, not a benchmark run -- I'll take the correction. But even as a ceiling it says the mechanism has real teeth. Here's what actually bugs me more: that r equals 0.99 judge number. That's two language models agreeing with each other, not with a human. On multi-span reasoning, where the accuracy gaps are widest, I'd want that checked against human labels before trusting the drop numbers at face value. 25 00:11:34,951 --> 00:12:02,769 [Dr. Ada Shannon] Right instinct, and there's a second loose thread next to it: that 15 to 35 percent more decode steps isn't free, and up to six percent of Gemma-4-12B's DA responses never terminate inside the fixed eight-thousand-token budget, which inflates both the failure count and the attended-token totals for that model. Before anyone quotes 52 percent as a clean number, it's worth asking how much of it sits on top of a budget cap that was fixed somewhat arbitrarily. 26 00:12:02,769 --> 00:12:44,007 [Hal Turing] Right, and that overhang doesn't go away by itself — fifteen to thirty-five percent more decode steps is structural to how DA reasons, not a knob you tune down. So if six percent of Gemma-4-12B's responses are hitting the eight-thousand-token wall, some slice of those fifty-two and thirty-one percent savings numbers are coming from runs that technically didn't finish. Would the accuracy drop and the token savings look the same at sixteen or thirty-two K? We don't know, because it's never tested. And eight thousand tokens as a hard ceiling feels almost quaint for anything agentic — one tool-heavy turn can chew through that on its own before the model's even done navigating. 27 00:12:44,007 --> 00:13:30,076 [Dr. Ada Shannon] That compounds with something that bugged me through the results section: the whole accuracy comparison rests on an LLM judge, Qwen-3.5-4B, validated only by agreeing with Gemini-3.1-Pro at Pearson r equals 0.99. That's two models agreeing with each other, not with a human. Fine for single-span retrieval, where the passkey's either in the answer or it isn't. But DA's accuracy gaps concentrate in multi-span reasoning, dialogue history, long-dependency QA — exactly the free-form answers where two judges can converge on the same wrong rubric reading and a correlation number would never catch it. Nobody spot-checked this against a human annotator, right where the paper's own numbers say DA hurts most. 28 00:13:30,076 --> 00:14:00,772 [Hal Turing] Oh — wait, sorry to cut you off, but that connects to something that worries me even more: they disabled thinking mode for every result in this paper, because their own preliminary experiments showed models fail to follow the DA protocol inside thinking traces. Every frontier long-context deployment running today leans on interleaved thinking — Claude, the Kimi and DeepSeek lines their own related-work section cites — so the validated benefit here sits on exactly the axis production systems are moving away from. 29 00:14:00,772 --> 00:14:40,850 [Dr. Ada Shannon] And that's where I start wondering how much of this is genuinely new versus a scaled rerun of something we already had. Tian Jin and coauthors published Self-Selected Attention Span in 2024 — same core idea, a model declaring its own attention span so the engine masks around it. But that work trained a separate model per task with hand-designed annotation syntax at two-thousand-token contexts. This paper's pitch is one fixed zero-shot protocol generalizing task-agnostic out to a hundred K, no training at all. Real jump in scale — but a jump on top of an existing mechanism, not a new one, and I think the framing sells it as more novel than it is. 30 00:14:40,850 --> 00:15:16,469 [Hal Turing] I actually disagree with you there, Ada. Going from a trained, per-task, two-K-token behavior to something an off-the-shelf model does zero-shot at fifty times the context length isn't incremental scaling — that's the whole finding. Jin's result shows the capability can be trained in; this result shows it's already latent and comes out through prompting alone, on models that never saw this objective during training. Those are categorically different claims, and only one of them says something new about what pretraining already encodes. 31 00:15:16,469 --> 00:15:39,411 [Dr. Ada Shannon] Sure, but 'zero-shot elicitation' and 'already latent' aren't the same claim either — you can't rule out that this chunked tool-call format is just a more effective annotation syntax than Jin's, riding on a few extra years of pretraining scale to make it stick. There's no head-to-head run against Jin's method on the same benchmarks here, so 'it was latent all along' is an interpretation of the result, not something they actually measured. 32 00:15:39,411 --> 00:16:18,095 [Hal Turing] That's fair, I'll grant the missing head-to-head. Where I think we do agree is on the paper's own system-two framing — they cite Weston and Sukhbaatar's System 2 Attention, out of Meta, from 2023, as the lineage, but that method realizes selection by regenerating the entire input once before answering. DA never regenerates; it masks attention live, mid-generation, as reasoning unfolds. So however the Jin question shakes out, the actual mechanism — block-level masking driven by declarations as they're emitted — is a genuinely different move than either prior line. 33 00:16:18,095 --> 00:17:00,588 [Dr. Ada Shannon] Where this gets practically interesting is what it plugs into. Global mode still pays full attention cost and eats most of DA's attended tokens, so pairing it with a lightweight-scan method like Quest to cover exactly that navigation phase is the obvious next step — the paper says as much itself. Speculative decoding stacks too, since DA's mask stays fixed within a mode, so drafting runs under one mask and only resets at the infrequent transitions, offsetting some of the step overhead we started with. But the idea I like best is reversible KV offloading: because DA never edits the sequence the way compaction does, an out-of-focus segment's cache can spill to host memory and get prefetched back the moment a focus tag names it again. 34 00:17:00,588 --> 00:17:47,120 [Hal Turing] That's the direction I'd bet real deployment on — post-training or RL on the protocol itself, since every number here is a zero-shot floor from a model that never saw this objective in training. Pair that with natural agentic segments, tool calls, turns, retrieved passages, instead of arbitrary two-K chunks cut into a static benchmark, and let the model survey a cheap in-context index instead of paying full cost every time it goes global. So where does that leave us? This paper turns attention selection from a pattern buried inside the network into text anyone can read and an engine can act on, at a cost close to vanilla — but everything today is a lower bound, not a ceiling. Thanks for listening.