A visual explainer of why million-token decoding is dominated by KV-cache traffic, how RetrievalAttention mixes static GPU retention with dynamic CPU-side ANN lookup, and where the method shines or risks failure.
Long-context decoding becomes an I/O problem: each new token rereads a growing KV history.
Static retention covers “always useful” tokens; dynamic retrieval handles query-specific needles.
Attention is often dynamically sparse: different queries light up different slivers of history.
RetrievalAttention sits in a crowded design space with compression, offloading, and heuristic sparse attention methods.
Additional arXiv IDs found in transcript: 2409.10516