podcast companion systems / RAG serving visual-first

CacheBlend for
Fast RAG Serving

A visual companion for the episode on reusing KV caches for repeated RAG chunks, where the core challenge is not exact prefix matching but preserving cross-chunk attention without paying full prefill cost.
arXiv 2405.16444 EuroSys 2025 / arXiv 2024→2025 Focus: TTFT, throughput, cache reuse

From query to first token

RAG makes prefill expensive because query + multiple retrieved chunks must be processed before generation starts. CacheBlend inserts chunk-level cache lookup, selective token repair, and KV fusion into that path.

Hover nodes to inspect where latency accumulates. The optimized path shifts work from full prompt prefill toward cache fetch + partial repair.
Repeated chunks are often not prompt prefixes. That is why exact prefix caching leaves obvious reuse opportunities on the table.

Why naive non-prefix reuse breaks

A reused chunk’s original KV state was computed under old left-context. In a new prompt, its tokens should attend to the query and earlier chunks. CacheBlend repairs only selected tokens whose cross-chunk dependencies are likely to matter.

cold / weak attention medium interaction hot / strong interaction repaired token rows
Selective recompute is layer-wise and request-dependent. The idea is surgical patching, not redoing the whole chunk.

Latency and throughput regimes

Mocked values follow the paper’s qualitative shape: full recompute is quality-safe but slow, prefix cache helps only narrow cases, naive non-prefix reuse is fastest but risks quality loss, and CacheBlend aims for a middle point with most of the speed and near-full quality.

The useful systems story is relative: CacheBlend mostly matters when repeated non-prefix chunks are common and prefill dominates user-visible latency.
The paper’s central empirical claim is that small update fractions can recover answer quality close to full recompute while preserving speed.

Where CacheBlend helps most

This section visualizes the episode’s caveat: the method shines in high-hit-rate, prefill-heavy RAG with repeated corpora, but the payoff shrinks if workloads are decode-heavy, highly dynamic, or already dominated by exact prefix sharing.

Axes: repeated non-prefix chunk rate × cross-chunk reasoning need. Best region for CacheBlend is upper-right-ish but not universal.
CacheBlend complements, rather than replaces, prefix sharing, paging, offloading, and fast attention kernels.

References

CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
Jiayi Yao et al., 2024/2025
paper
arXiv: 2405.16444
Prompt Cache: Modular Attention Reuse for Low-Latency Inference
Yao Fu et al., 2024
scholar
CacheGen
KV cache compression and streaming for fast serving
scholar
RadixAttention
Exact prefix-oriented KV sharing in serving stacks
scholar
vLLM / PagedAttention
Practical inference engine context for implementation
scholar
Related podcast episodes
Prefix cache, CacheSlide, cross-datacenter KV, KVSwap, FengHuang, speculative decoding
podcast archive