This page treats the paper as a serving problem: how do you reuse retrieved context without erasing the cross-chunk interactions that multi-hop question answering depends on? Explore the prompt pipeline, attention heatmaps, quality-latency frontiers, and the offline-plus-online design that makes FusionRAG Cache different from plain chunk reuse.
Retrieval improves grounding, but every retrieved chunk extends the prefill path before the model can emit the first answer token. The central question is not whether caching helps, but whether chunk-level reuse can preserve the interactions that the full assembled prompt would have created.
Mock numbers below mirror the paper’s framing: the biggest win is shrinking prefill work while keeping enough cross-document signal to avoid quality collapse.
Classic prefix caching only works when requests share one contiguous prompt prefix. RAG rarely does. Retrieved passages recur across requests, but their order and neighbors change, so a chunk encoded in isolation is not the same chunk encoded inside the final evidence pack.
The heatmaps show why naive chunk reuse can fail. Tokens inside one passage are easy to cache, but the bridge between two passages often carries the answer. FusionRAG tries to preserve more of that structure through offline enrichment and online token-focused repair.
Rows are query-side salient tokens and columns are retrieved evidence tokens. The diagonal blocks represent within-chunk attention. The off-diagonal bridges represent multi-document reasoning, citation stitching, or disambiguation across separate passages.
When chunks are cached independently, the model may see each paragraph as an island. That lowers compute, but it can also erase the very cross-chunk signal needed to combine evidence from chunk B with definitions, entities, or caveats introduced in chunk A.
The paper’s claim is strongest in the low-recompute regime: spend a small budget on online repair and keep much more quality than naive chunk reuse. The frontier below uses realistic mock data to mirror that shape across QA-heavy RAG workloads.
Left side means tiny recompute budgets and aggressive reuse. Right side approaches full recomputation. A strong RAG serving method moves upward without collapsing leftward speedups.
FusionRAG’s conceptual move is to stop treating cache reuse as all-or-nothing. The system pre-bakes some cross-document context ahead of time, then spends a small live-request budget on the tokens that appear most worth repairing.
This kind of method is strongest when prompts are long, retrieved corpora are fairly stable, and TTFT is the main user-visible pain. It is weaker when corpora churn fast, access control is highly personalized, or bottlenecks sit elsewhere in the pipeline.
Offline fusion is not free. It can increase storage, complicate invalidation when documents change, and make cache management harder in multi-tenant systems. The speedup story also does not solve retrieval stalls, reranker cost, or decoder-heavy workloads by itself.
Compact source list for the episode and its neighboring cache-reuse literature.