KVSwap for Disk-Aware
Long-Context On-Device Inference

A visual companion for the episode on disk-backed KV cache offloading: why server-style GPU→CPU offload breaks on unified-memory devices, how KVSwap predicts which history matters, and how prefetch + buffering + storage-shaped reads make NVMe / UFS / eMMC barely fast enough to sustain long-context decoding.

arXiv2511.11907 TopicKV cache + storage systems Core claimfull KV on disk, tiny guidance in RAM Contextsup to 32K
same-memory throughput uplift
1.8×
NVMe vs naive disk offload
same-memory throughput uplift
4.1×
eMMC vs naive disk offload
KV memory footprint
~11× less
than an in-RAM serving setup comparison
critical constraints
BW · Lat · RA
bandwidth, latency, read amplification

Decode Pipeline: from query to disk-guided exact attention

KVSwap keeps the full-fidelity KV cache on storage, but preserves a compact in-memory key-side sketch to predict which groups will be needed next. Hover nodes and arrows.

memory tier predictor / compact K storage / disk active data path

Memory-tier pressure map

Unified memory means “GPU to CPU offload” often stays inside the same scarce RAM pool. KVSwap drops to storage instead.

Per-token timeline

In KVSwap, I/O is issued ahead of compute and partially hidden under decoding work.

Storage physics: why tiny random reads are poison

Disk offloading only works if accesses are reshaped around flash behavior. Toggle between media and compare logical requests vs effective physical transfer.

Read amplification matrix

Rows = request granularity; columns = access locality. Cooler is better. Hover each cell.

Sequentialized group reads vs scattered token reads

KVSwap predicts useful groups, reads them in chunk-friendly order, then reuses them from an in-memory buffer across nearby steps.

Selection heatmaps: compact K sketch → choose groups → fetch exact KV

The predictor should be cheap enough to stay in RAM but accurate enough to avoid missing relevant old context. Use the step buttons.

History score heatmap

Mock attention-like relevance over token groups across layers. Hover to inspect old spans and “important but inconvenient” regions.

Predictor accuracy / compression tradeoff

More aggressive compression lowers in-RAM footprint, but selection accuracy degrades and extra reloads erase gains.

Performance envelope under tight memory budgets

Illustrative mock data based on the episode’s reported trends: stronger gains on slower media because naive patterns become more pathological there.

Comparison bars

Methods: in-RAM serving style, naive disk offload, and KVSwap. Hardware/media profiles on the x-axis.

Scaling with context length

As context grows, disk traffic dominates unless selection + overlap keep the active working set small and predictable.

References

KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
Zhang, Xia, Wang · 2025 · arXiv:2511.11907
FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Chen et al. · 2023 · tiered offloading baseline
vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention
Kwon et al. · 2023 · paged KV management
RetrievalAttention
Shen et al. · 2024 · selective long-context retrieval
SnapKV, H2O, StreamingLLM, PyramidInfer, KeyDiff, CHESS
Compression / eviction / selective retention family
Episode note
No additional transcript arXiv IDs detected beyond 2511.11907.