AI Post Transformers · Episode Companion

Fast State Restoration for Evicted LLM KV Caches

arXiv:2410.05004 ↗ Shiwei Gao, Youmin Chen, Jiwu Shu — Tsinghua University EuroSys 2025 · Rotterdam Submitted Oct 7, 2024

When GPU memory pressure evicts a conversation's KV cache, today's serving systems either recompute it (20–26× slower) or stream it back from storage (6.5–13× slower). HCache proposes a third path: cache the hidden state one layer upstream, and rebuild K and V on demand with a cheap matrix multiply.

Two Bad Options

If an evicted conversation's state isn't cached KV-for-KV, there are exactly two fallbacks today — and both are expensive relative to never having evicted anything at all.

Why Eviction Is the Common Case

A single A100-40GB holds ~17K tokens of KV cache for Llama2-13B, or ~48K for Llama2-7B. Toggle between the two workload traces to see how few conversations that actually buys.

Overhead Grows With Context Length

Illustrative, normalized severity of restoration overhead across context-length buckets. Hover a cell — recompute's quadratic attention cost is what makes long L-Eval-style contexts hurt the most.

Three Restoration Paths

Same evicted layer, three ways to get K and V back. Select a path to see its relative I/O and compute cost below.

Pipelining Transfer and Compute

Neither the PCIe link nor the GPU should sit idle. While layer i's hidden state streams in, layer i-1's K/V projection runs concurrently on the GPU.

Layer-wise vs Token-wise Splitting

The bubble-free scheduler splits layers between HCache restoration and a complementary method to keep transfer and compute balanced. Token-wise splitting looked more flexible on paper — it broke cuBLAS's tuned GEMM shapes instead.

Chunked, Striped Storage

Hidden states are split into fixed 64-token chunks and striped round-robin across every SSD, so restoring one layer pulls from all drives in parallel.

Decoding barely notices HCache running in the background — under 4% TBT overhead — because a background CPU daemon flushes chunks to SSD off the critical path, instead of stalling generation with a direct write.

MHA vs GQA: Does the 2× I/O Win Survive?

All three test models (Llama2-7B, Llama2-13B, OPT-30B) use plain multi-head attention. The hidden state HCache caches is captured before the K/V projection, so it stays full-width no matter how narrow that projection gets under grouped-query attention.

Two More Asterisks

Section 7 calls the GQA case "beyond scope" — but K = Wk·H doesn't care how narrow Wk is, so this is an economics gap, not a structural one. Two more comparisons tilt the reported numbers further in HCache's favor.

Handicapped baseline: the KV-offload baseline is a reimplementation of AttentionStore (CachedAttention, Alibaba, 2024) that explicitly skips its decoupled position-embedding optimization — never open-sourced, so HCache is compared against a slower rebuild of its rival.
Missing comparison: CacheGen (Liu et al., 2024) pairs KV offload with quantization-based compression. HCache never runs this comparison — quantizing hidden states is future work, not a tested result.

References

  1. Fast State Restoration in LLM Serving with HCache — Shiwei Gao, Youmin Chen, Jiwu Shu, 2024. arxiv.org/abs/2410.05004
  2. Training Deep Nets with Sublinear Memory Cost — Tianqi Chen, Bin Xu, Chiyuan Zhang, Carlos Guestrin, 2016. Google Scholar ↗
  3. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, et al., 2023. Google Scholar ↗
  4. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, et al., 2024/2025. Google Scholar ↗
  5. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Yuhan Liu, Junchen Jiang, et al., 2024. Google Scholar ↗
  6. CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion — Jiayi Yao, Hanchen Li, Yuhan Liu, et al., 2025. Google Scholar ↗
  7. SGLang: Efficient Execution of Structured Language Model Programs — Lianmin Zheng, Liangsheng Yin, Ying Sheng, et al., 2024. Google Scholar ↗
  8. Cost-Efficient LLM Serving for Multi-turn Conversations with CachedAttention — Bin Gao, Zhuomin He, Puru Sharma, et al., 2024 (USENIX ATC). Google Scholar ↗
  9. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie, James Lee-Thorp, Michiel de Jong, et al., 2023 (EMNLP). Google Scholar ↗
  10. Prompt Cache: Modular Attention Reuse for Low-Latency Inference — In Gim, Guojun Chen, Seung-seob Lee, et al., 2024 (MLSys). Google Scholar ↗
  11. Efficiently Programming Large Language Models using SGLang — Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, et al., 2023/2024. Google Scholar ↗
  12. FlexGen: High-Throughput Generative Inference of LLMs with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, et al., 2023 (ICML). Google Scholar ↗