This companion page visualizes Re-Prefill: reload a shared prefix KV cache from CPU/SSD, compute a request-specific suffix, then emit the first token. The paper argues the true bottlenecks are not transformer FLOPs alone, but read amplification and serialized I/O→compute dependency. The visuals below show how granularity-aligned contiguous chunks, speculative prefetching, and attention-guided residency try to reduce those stalls.
Switch between a naïve offload path and the paper’s co-designed path. The goal is to turn “decide → read → compute” from a single-file queue into an overlapped pipeline.
Hover the heatmaps. The left panel shows fine semantic importance scattered across storage blocks. The right panel reorganizes the same useful regions into contiguous chunk units used for selection, placement, transfer, and residency.
Step through layers to see how similar important chunk indices across nearby layers let the system prefetch ahead. The timeline contrasts serialized execution with overlapped I/O and compute.
The chart uses realistic mock values consistent with the discussion. Toggle the comparison lens: raw Re-Prefill time, effective bytes read, or a simple quality-latency Pareto sketch.
The left visual shows GPU/CPU/SSD residency updated by attention. The right visual situates ContiguousKV among compression, paging, and retention systems discussed in the episode.
Compact source list for the episode and nearby systems it compares against.