A visual companion for the episode on disk-backed KV cache offloading: why server-style GPU→CPU offload breaks on unified-memory devices, how KVSwap predicts which history matters, and how prefetch + buffering + storage-shaped reads make NVMe / UFS / eMMC barely fast enough to sustain long-context decoding.
KVSwap keeps the full-fidelity KV cache on storage, but preserves a compact in-memory key-side sketch to predict which groups will be needed next. Hover nodes and arrows.
Disk offloading only works if accesses are reshaped around flash behavior. Toggle between media and compare logical requests vs effective physical transfer.
The predictor should be cheap enough to stay in RAM but accurate enough to avoid missing relevant old context. Use the step buttons.
Illustrative mock data based on the episode’s reported trends: stronger gains on slower media because naive patterns become more pathological there.