Interactive Visualization · GPU Scheduling · OSDI 2025 thread

NanoFlow and the Future of LLM Serving

A visual map of the shift from memory-only explanations toward overlap-aware serving: prefill, decode, GEMMs, attention, and communication stop marching in sequence and start sharing the machine.

Reported Gain
1.91×
Throughput lift versus major serving baselines in the paper’s favorable mixed-work regime.
Search Space
4D
Nano-batch count, size, operation order, and GPU resource split are chosen jointly.
Key Tension
Overlap
Local duplication can be worth it when idle tensor-core time is the real system tax.
Counterpoint
Memory
Long-context and smaller-batch regimes can swing the bottleneck back toward cache traffic.

From Batch Thinking to Nano-Operations

Each step below shrinks the scheduling unit. The visual moves from static batching to chunked prefill and then to NanoFlow’s fine-grained overlap of compute, memory, and communication work.

Compute-heavy GEMM / FFN Memory-heavy attention / KV traffic Communication / collectives Scheduling boundary or phase split

Overlap Lab

Switch serving policies to see how the same prompt+decode workload occupies a GPU. The goal is not less work everywhere; it is less waiting between unlike kernels.

Throughput Frontier

Mock data mirrors the paper’s narrative: big gains in mixed, tensor-parallel workloads; tighter margins when memory pressure or smaller online batches dominate.

vLLM / strong baseline Chunked-prefill family Phase-split family NanoFlow searched overlap

Bottleneck Atlas

Heatmaps show why “memory-bound” and “compute-bound” can both be true. Hover cells to inspect intensity, overlap opportunity, and interference risk across the serving stack.

High overlap potential appears where one operation is compute-dense and the other is memory- or communication-skewed. High interference appears when both fight for the same resource at the same moment.

References

Compact source map for the visual thread. The transcript did not contain any extra arXiv-style IDs beyond the three already provided: 2601.11822, 2410.18038, 1811.06965.

NanoFlow: Towards Optimal Large Language Model Serving Throughput OSDI 2025 paper PDF
Chunked-prefill / prefill-decode mixing thread arXiv:2601.11822 · arXiv:2410.18038 2601.11822 2410.18038
FlashAttention and IO-aware attention thinking Attention as data movement scholar link
PagedAttention / vLLM memory substrate KV allocation and fragmentation scholar link
Orca and iteration-level scheduling Serving as scheduler design scholar link
Sarathi-Serve and throughput-latency tradeoff Chunked prefill frontier paper PDF
DistServe and disaggregated serving Separate prefill from decode scholar link
Foundational transformer reference arXiv:1811.06965 1811.06965