Agent-based LLM systems have a modular prompt structure: fixed system instructions, dynamic content (API results, tool outputs), and conversation history. When content shifts position, traditional KV caching breaks down.
When cached KVs are reused at different positions, their positional encodings become stale. This visualization shows how position mismatch creates attention drift.
CacheSlide introduces RPDC (Relative-Position-Dependent Caching) as a middle ground between rigid PDC and expensive PIC.
Reuses KVs only at exact absolute positions. Fast but inflexible—any insertion breaks reuse for downstream segments.
RPDC achieves 85-90% reuse with minimal correction overhead, while PIC requires full attention recomputation.
CCPE replaces absolute position indices with chunk-relative offsets, making cached KVs portable across requests.
Blends cached KVs with selectively recomputed correction tokens using learned per-layer weights.
Blending Formula: final_KV = α × cached_KV + (1-α) × correction_KV
Alpha weights are calibrated offline per layer, then fixed at inference time.
Only 5-10% of tokens in cached segments are recomputed to correct positional drift.
CacheSlide reduces prefill latency by 3.11-4.3× compared to vanilla vLLM on ToolBench and AgentBench.
Throughput improvements scale with batch size, achieving 5.8× gains on LLaMA-7B at high concurrency.
Task success rate remains above 94% across all tested models with correction ratio at 10%.