Prompt-Time Attention Is Not Decode-Time Need
Hover the matrices. Left: broad prefill salience. Right: decode-conditioned need after answer generation begins. The sharper the divergence, the weaker prompt-only eviction signals become.
Token Retention Map
Each lane is a prompt segment. The active mode changes which spans survive under the same KV budget.
How LAQ Replaces a Passive Window with an Active Probe
Step through the decode preparation path. The draft tokens are not mainly for acceptance. They act as a query probe to estimate which cached prefix tokens future generation will revisit.
Head-Level Utility Heatmap
Mock head-by-token demand under decode. Hover cells to inspect head specialization. Brighter cells mark likely retrieval-heavy interactions that eviction must avoid deleting.
Budget Sensitivity Under Decode Compression
Move the slider. This mock chart keeps the episode’s framing: gains are most visible in tight-budget decode regimes, and the advantage narrows as memory pressure relaxes.
Deployment Scoreboard
The paper’s idea is strong on eviction consistency. It is weaker on end-to-end serving evidence. The chart separates benchmark lift from infra trust and stack fit.
LAQ Is a Narrow Tool Inside a Larger Inference Stack
This map separates bottlenecks. Speculative decoding reduces token-generation waste. GQA and architecture changes reduce KV size structurally. LAQ targets decode-stage cache selection.
References
Compact source trail for the papers and podcast episodes most directly discussed here.