CacheSlide: Position-Aware KV Cache Reuse for Agent LLMs

Yang Liu et al. | Shanghai Jiao Tong University, Inspur, Peking University, Huawei Cloud
USENIX FAST'26 | 2026

View on arXiv

The Agent Workflow Challenge

Agent-based LLM systems have a modular prompt structure: fixed system instructions, dynamic content (API results, tool outputs), and conversation history. When content shifts position, traditional KV caching breaks down.

Latency Reduction
3.11-4.3×
Throughput Gain
3.5-5.8×
Accuracy Loss
<1%
Correction Tokens
5-10%

PMKD: Positionally Misaligned KV Drift

When cached KVs are reused at different positions, their positional encodings become stale. This visualization shows how position mismatch creates attention drift.

Low Drift (0.0-0.3)
Medium Drift (0.3-0.7)
High Drift (0.7-1.0)

Three Caching Paradigms

CacheSlide introduces RPDC (Relative-Position-Dependent Caching) as a middle ground between rigid PDC and expensive PIC.

Position-Dependent Caching (PDC)

Reuses KVs only at exact absolute positions. Fast but inflexible—any insertion breaks reuse for downstream segments.

Computational Cost Comparison

RPDC achieves 85-90% reuse with minimal correction overhead, while PIC requires full attention recomputation.

Chunked Contextual Position Encoding (CCPE)

CCPE replaces absolute position indices with chunk-relative offsets, making cached KVs portable across requests.

Weighted Correction Attention

Blends cached KVs with selectively recomputed correction tokens using learned per-layer weights.

Blending Formula: final_KV = α × cached_KV + (1-α) × correction_KV

Alpha weights are calibrated offline per layer, then fixed at inference time.

Correction Token Selection Strategy

Only 5-10% of tokens in cached segments are recomputed to correct positional drift.

Latency: Time-to-First-Token (TTFT)

CacheSlide reduces prefill latency by 3.11-4.3× compared to vanilla vLLM on ToolBench and AgentBench.

Throughput: Requests per Second

Throughput improvements scale with batch size, achieving 5.8× gains on LLaMA-7B at high concurrency.

Accuracy vs. Speed Trade-off

Task success rate remains above 94% across all tested models with correction ratio at 10%.

Key References

  • Vaswani et al. (2017) — Attention Is All You Need (original Transformer) arXiv:1706.03762
  • Su et al. (2021) — RoFormer: Enhanced Transformer with Rotary Position Embedding (RoPE) arXiv:2104.09864
  • Touvron et al. (2023) — LLaMA: Open and Efficient Foundation Language Models arXiv:2302.13971
  • Kwon et al. (2023) — Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM) arXiv:2309.06180
  • Yao et al. (2024) — CacheBlend: Fast Large Language Model Serving with Cached Knowledge Fusion arXiv:2405.16444
  • Qin et al. (2023) — ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs arXiv:2307.16789
  • Liu et al. (2023) — AgentBench: Evaluating LLMs as Agents arXiv:2308.03688
  • Yao et al. (2022) — ReAct: Synergizing Reasoning and Acting in Language Models arXiv:2210.03629