Interactive podcast companion arXiv: 2409.10516 2024 paper • long-context inference • vector retrieval

RetrievalAttention for Long-Context LLM Inference

A visual explainer of why million-token decoding is dominated by KV-cache traffic, how RetrievalAttention mixes static GPU retention with dynamic CPU-side ANN lookup, and where the method shines or risks failure.

125 GB
KV cache / 1M tokens
for Llama-3-8B
32.8s → 1765s
decode latency growth
128K to 1M context
1–3%
token access fraction
claimed by system
0.188 s/token
flagship reported setup
128K on RTX 4090

What to look at

Use the tabs to switch between four visual views: bottlenecks, retrieval mechanics, attention heatmaps, and system tradeoffs.

Core idea

Don’t scan the entire KV history on every generated token. Keep a tiny always-on set on GPU, push the rest to CPU, retrieve likely relevant KV entries with ANN, then compute attention on only that subset.

Key caveat

Attention lookup is not generic vector search: query and key distributions are mismatched, so vanilla ANN can need to inspect 30–50% of keys to preserve quality.

Serving pain moves from FLOPs to memory traffic

Long-context decoding becomes an I/O problem: each new token rereads a growing KV history.

compute / active GPU work memory-resident state transfer / lookup overhead dominant bottleneck
Mock reconstruction from episode figures: decoding latency grows steeply with context, while KV-cache size scales roughly linearly.

Hybrid memory pipeline

Static retention covers “always useful” tokens; dynamic retrieval handles query-specific needles.

Interactive stepper: show how the query flows through retention, CPU ANN, reranking, and sparse attention assembly.
The paper’s key warning: attention queries are out-of-distribution relative to indexed keys, so generic ANN geometry can fail.

Dense attention vs retrieval-selected attention

Attention is often dynamically sparse: different queries light up different slivers of history.

Rows are generated queries; columns are historical tokens. Hover cells to inspect local attention mass.
Head-by-head mock taxonomy: some heads look streaming/local, others retrieval-heavy, suggesting hybrids like DuoAttention-style splits.

Performance landscape

RetrievalAttention sits in a crowded design space with compression, offloading, and heuristic sparse attention methods.

Relative mock comparison under equalized long-context serving assumptions, based on the qualitative framing from the episode.
Latency-quality-memory tradeoff map. Bubble size ≈ implementation complexity / systems overhead.

References

Additional arXiv IDs found in transcript: 2409.10516

Listen to the Episode
podcast.do-not-panic.com