Language Models Learn to Fetch Only What They Need

AI Post Transformers · KAIST AI · Declarative Attention · KV-cache sparsity
arXiv:2609.02737 ↗

The paper "Language Models Can Control Their Own Attention" tackles the long-context memory bottleneck: at a million tokens, generating each token drags ~15GB of KV cache through memory — comparable to reloading the model's entire active parameter set. Instead of an external scorer guessing what matters, the model declares it, in its own reasoning, and the inference engine skips loading the rest.

The bill: 15GB per token at a million-token context

Every decoded token means re-reading the entire KV cache. At long context this read dominates, not the matmuls. Bars below compare that per-token memory traffic against the cost of reloading the model's own active parameters.

KV cache read per token (1M ctx) Reloading active params (Qwen-3.5-397B, 17B active)

Two ways to decide what to skip

Toggle between the extrinsic-scoring lineage (StreamingLLM → H2O → Quest) and this paper's intrinsic alternative. Extrinsic methods still require a pass over context statistics before anything can be discarded; Declarative Attention lets the model say it directly.

Worked example: founding year → IPO year → subtract

The model cycles global (skim everything), focus (read one chunk), and local (attend only to its own scratch work) modes. Click a step.

Magic chunks: fake tool-call theater

Context is split into ~2000-token chunks and presented as a fabricated tool-call transcript, so chunk boundaries land on turn-delimiter tokens the model already tracks from post-training. No tool actually runs.

Attention footprint by decode step

Hover a cell. Rows are decode steps in the worked example above; columns are magic chunks. Color = fraction of that chunk's KV blocks actually loaded.

0.0 skipped 0.5 1.0 fully loaded

Attended tokens vs. accuracy drop, six models

Switch arms. DA-no-mask is the control: same chunked-reasoning prompt, mask disabled. If savings came from format alone rather than masking, this arm would look like DA — it doesn't.

Roofline wall-clock ratio (decode only)

DA / vanilla decode time at each op's own hardware ceiling (40% MFU, 70% MBU, single B200). A theoretical floor, not a stopwatch benchmark — and it excludes prefill.

Overhead that doesn't disappear

DA reasoning costs more decode steps. On the smallest model, some of that overhead hits a fixed 8K-token budget cap and never finishes.

Sparse-attention lineage

Every prior method still pays for a full pass over context statistics before discarding anything. Declarative Attention is the first to make the decision live, mid-generation, from the model's own text.

MethodYearDecides viaFull context scan first?Training needed?Live masking mid-gen?
StreamingLLM (sinks)2023Position heuristicNoNoNo
H2O (heavy hitters)2023Cumulative attn scoreYesNoNo
Quest2024Query-aware page scoreYesNoNo
Self-Selected Attn Span2024Trained span headNoYes, per-taskYes
System 2 Attention2023Full input regenerationYesNoNo
Declarative Attention2026Model's own reasoning traceNoZero-shotYes

Where the trust breaks down: judge reliability by task type

The LLM judge (Qwen-3.5-4B vs Gemini-3.1-Pro, r=0.99) agrees best on tasks with a single verifiable answer. Its two models agreeing with each other is not the same as agreeing with a human — and DA's accuracy gaps concentrate exactly where judge confidence is weakest.

Multi-span reasoning and dialogue history — where DA hurts accuracy most — are also where two LLM judges are most likely to converge on the same wrong rubric reading. No human spot-check was run on these cases.

Self-report vs. independent estimate

The hosts' core disagreement: is the model's self-declared relevance categorically riskier than an external scorer's guess?

Case for intrinsic

A model competent enough to answer a question already implicitly knows which part of the document matters — that's what reasoning over a document is. Zero-shot, no training, generalizes past 100K tokens on models that never saw this objective.

Case for caution

An extrinsic scorer is wrong sometimes but is at least independent. Here the same system grades its own homework — an error doesn't show up as recoverable noise, it becomes an actual hole in what the model can see.

Sources cited in this episode