When GPU memory pressure evicts a conversation's KV cache, today's serving systems either recompute it (20–26× slower) or stream it back from storage (6.5–13× slower). HCache proposes a third path: cache the hidden state one layer upstream, and rebuild K and V on demand with a cheap matrix multiply.
If an evicted conversation's state isn't cached KV-for-KV, there are exactly two fallbacks today — and both are expensive relative to never having evicted anything at all.
A single A100-40GB holds ~17K tokens of KV cache for Llama2-13B, or ~48K for Llama2-7B. Toggle between the two workload traces to see how few conversations that actually buys.
Illustrative, normalized severity of restoration overhead across context-length buckets. Hover a cell — recompute's quadratic attention cost is what makes long L-Eval-style contexts hurt the most.
Same evicted layer, three ways to get K and V back. Select a path to see its relative I/O and compute cost below.
Neither the PCIe link nor the GPU should sit idle. While layer i's hidden state streams in, layer i-1's K/V projection runs concurrently on the GPU.
The bubble-free scheduler splits layers between HCache restoration and a complementary method to keep transfer and compute balanced. Token-wise splitting looked more flexible on paper — it broke cuBLAS's tuned GEMM shapes instead.
Hidden states are split into fixed 64-token chunks and striped round-robin across every SSD, so restoring one layer pulls from all drives in parallel.
Decoding barely notices HCache running in the background — under 4% TBT overhead — because a background CPU daemon flushes chunks to SSD off the critical path, instead of stalling generation with a direct write.
All three test models (Llama2-7B, Llama2-13B, OPT-30B) use plain multi-head attention. The hidden state HCache caches is captured before the K/V projection, so it stays full-width no matter how narrow that projection gets under grouped-query attention.
Section 7 calls the GQA case "beyond scope" — but K = Wk·H doesn't care how narrow Wk is, so this is an economics gap, not a structural one. Two more comparisons tilt the reported numbers further in HCache's favor.