The paper "Language Models Can Control Their Own Attention" tackles the long-context memory bottleneck: at a million tokens, generating each token drags ~15GB of KV cache through memory — comparable to reloading the model's entire active parameter set. Instead of an external scorer guessing what matters, the model declares it, in its own reasoning, and the inference engine skips loading the rest.
Every decoded token means re-reading the entire KV cache. At long context this read dominates, not the matmuls. Bars below compare that per-token memory traffic against the cost of reloading the model's own active parameters.
Toggle between the extrinsic-scoring lineage (StreamingLLM → H2O → Quest) and this paper's intrinsic alternative. Extrinsic methods still require a pass over context statistics before anything can be discarded; Declarative Attention lets the model say it directly.
The model cycles global (skim everything), focus (read one chunk), and local (attend only to its own scratch work) modes. Click a step.
Context is split into ~2000-token chunks and presented as a fabricated tool-call transcript, so chunk boundaries land on turn-delimiter tokens the model already tracks from post-training. No tool actually runs.
Hover a cell. Rows are decode steps in the worked example above; columns are magic chunks. Color = fraction of that chunk's KV blocks actually loaded.
Switch arms. DA-no-mask is the control: same chunked-reasoning prompt, mask disabled. If savings came from format alone rather than masking, this arm would look like DA — it doesn't.
DA / vanilla decode time at each op's own hardware ceiling (40% MFU, 70% MBU, single B200). A theoretical floor, not a stopwatch benchmark — and it excludes prefill.
DA reasoning costs more decode steps. On the smallest model, some of that overhead hits a fixed 8K-token budget cap and never finishes.
Every prior method still pays for a full pass over context statistics before discarding anything. Declarative Attention is the first to make the decision live, mid-generation, from the model's own text.
| Method | Year | Decides via | Full context scan first? | Training needed? | Live masking mid-gen? |
|---|---|---|---|---|---|
| StreamingLLM (sinks) | 2023 | Position heuristic | No | No | No |
| H2O (heavy hitters) | 2023 | Cumulative attn score | Yes | No | No |
| Quest | 2024 | Query-aware page score | Yes | No | No |
| Self-Selected Attn Span | 2024 | Trained span head | No | Yes, per-task | Yes |
| System 2 Attention | 2023 | Full input regeneration | Yes | No | No |
| Declarative Attention | 2026 | Model's own reasoning trace | No | Zero-shot | Yes |
The LLM judge (Qwen-3.5-4B vs Gemini-3.1-Pro, r=0.99) agrees best on tasks with a single verifiable answer. Its two models agreeing with each other is not the same as agreeing with a human — and DA's accuracy gaps concentrate exactly where judge confidence is weakest.
The hosts' core disagreement: is the model's self-declared relevance categorically riskier than an external scorer's guess?
A model competent enough to answer a question already implicitly knows which part of the document matters — that's what reasoning over a document is. Zero-shot, no training, generalizes past 100K tokens on models that never saw this objective.
An extrinsic scorer is wrong sometimes but is at least independent. Here the same system grades its own homework — an error doesn't show up as recoverable noise, it becomes an actual hole in what the model can see.