AI Post Transformers Systems + LLM Serving arXiv:2601.13631 Paper date: 2026-01-20

ContiguousKV for Faster
LLM Prefill KV Reuse

This companion page visualizes Re-Prefill: reload a shared prefix KV cache from CPU/SSD, compute a request-specific suffix, then emit the first token. The paper argues the true bottlenecks are not transformer FLOPs alone, but read amplification and serialized I/O→compute dependency. The visuals below show how granularity-aligned contiguous chunks, speculative prefetching, and attention-guided residency try to reduce those stalls.

Headline claim
3.85×
Re-Prefill speedup vs IMPRESS
Primary pain
I/O
semantic unit ≠ storage block
Workload fit
Shared
RAG, search, multi-turn QA

Re-Prefill as a storage-bound pipeline

Switch between a naïve offload path and the paper’s co-designed path. The goal is to turn “decide → read → compute” from a single-file queue into an overlapped pipeline.

Pipeline flow diagram

storage / SSD transfer / CPU-GPU path compute / suffix + attention stall / amplified wait

Stage latency composition

reload bytes
6.8 GB
useful bytes
1.9 GB
overlap ratio
12%
TTFT slice
480 ms
Baseline sketch: fine-grained semantic selection lands on coarse storage blocks, so each “useful” token region drags extra bytes and forces more waiting before compute can proceed.

Granularity mismatch → read amplification

Hover the heatmaps. The left panel shows fine semantic importance scattered across storage blocks. The right panel reorganizes the same useful regions into contiguous chunk units used for selection, placement, transfer, and residency.

Mismatched layout: semantic picks scattered across blocks

cold selected / warm hot / costly

Granularity-aligned ContiguousChunk layout

read amplification ≈ 3.6×

Speculative asynchronous prefetching

Step through layers to see how similar important chunk indices across nearby layers let the system prefetch ahead. The timeline contrasts serialized execution with overlapped I/O and compute.

Timeline: serialized vs overlapped

SSD read DMA / transfer layer compute stall bubble speculative prefetch

Cross-layer chunk similarity matrix

Mock similarity values illustrate the empirical claim used to justify prefetching: nearby layers often care about overlapping prefix regions.

Performance bars plus fairness questions

The chart uses realistic mock values consistent with the discussion. Toggle the comparison lens: raw Re-Prefill time, effective bytes read, or a simple quality-latency Pareto sketch.

Method comparison

What the skeptical listener asks for

stronger paper support higher uncertainty / needs ablation

Attention-guided residency and the neighboring design space

The left visual shows GPU/CPU/SSD residency updated by attention. The right visual situates ContiguousKV among compression, paging, and retention systems discussed in the episode.

Chunk residency by tier

Related-work map

References

Compact source list for the episode and nearby systems it compares against.

ContiguousKV
Jing Zou et al., 2026
arXiv:2601.13631
H2O
Zhenyu Zhang et al., 2023
Scholar link
Scissorhands
Yuhui Wang et al., 2023
Scholar link
SnapKV
Yao Fu et al., 2024
Scholar link
IMPRESS
Importance-based KV offloading, 2024
Scholar link
CacheGen
Liu et al., 2024
Scholar link
MiniCache
Liu et al., 2024
Scholar link
vLLM / PagedAttention
Kwon et al., 2023
Scholar link
StreamingLLM
Xiao et al., 2024
Scholar link
Podcast episode links
Prefix cache / SolidAttention / KVSwap / Prefill-as-a-Service / Speculative Decoding / Lookahead Q-Cache
podcast.do-not-panic.com