RAG makes prefill expensive because query + multiple retrieved chunks must be processed before generation starts. CacheBlend inserts chunk-level cache lookup, selective token repair, and KV fusion into that path.
A reused chunk’s original KV state was computed under old left-context. In a new prompt, its tokens should attend to the query and earlier chunks. CacheBlend repairs only selected tokens whose cross-chunk dependencies are likely to matter.
Mocked values follow the paper’s qualitative shape: full recompute is quality-safe but slow, prefix cache helps only narrow cases, naive non-prefix reuse is fastest but risks quality loss, and CacheBlend aims for a middle point with most of the speed and near-full quality.
This section visualizes the episode’s caveat: the method shines in high-hit-rate, prefill-heavy RAG with repeated corpora, but the payoff shrinks if workloads are decode-heavy, highly dynamic, or already dominated by exact prefix sharing.