A visual map of the shift from memory-only explanations toward overlap-aware serving: prefill, decode, GEMMs, attention, and communication stop marching in sequence and start sharing the machine.
Each step below shrinks the scheduling unit. The visual moves from static batching to chunked prefill and then to NanoFlow’s fine-grained overlap of compute, memory, and communication work.
Switch serving policies to see how the same prompt+decode workload occupies a GPU. The goal is not less work everywhere; it is less waiting between unlike kernels.
Mock data mirrors the paper’s narrative: big gains in mixed, tensor-parallel workloads; tighter margins when memory pressure or smaller online batches dominate.
Heatmaps show why “memory-bound” and “compute-bound” can both be true. Hover cells to inspect intensity, overlap opportunity, and interference risk across the serving stack.
High overlap potential appears where one operation is compute-dense and the other is memory- or communication-skewed. High interference appears when both fight for the same resource at the same moment.
Compact source map for the visual thread. The transcript did not contain any extra arXiv-style IDs beyond the three already provided: 2601.11822, 2410.18038, 1811.06965.
OSDI 2025
paper PDF
arXiv:2601.11822 · arXiv:2410.18038
2601.11822
2410.18038
Attention as data movement
scholar link
KV allocation and fragmentation
scholar link
Serving as scheduler design
scholar link
Chunked prefill frontier
paper PDF
Separate prefill from decode
scholar link
arXiv:1811.06965
1811.06965