A visual companion about long-context LLM serving where the bottleneck shifts from token generation to recovering old KV state fast enough to cut time-to-first-token.
Switch the mode to see how the same request hits different bottlenecks. The diagram emphasizes where waiting time accumulates before the first token appears.
Late tokens cost more to recompute because they drag a larger history window. Hover the matrix, then step the scheduler to watch the recompute/load boundary move under contention.
Use the hardware and traffic toggles. This chart is mock data shaped to match the episode’s claim: the scheduler looks best when bandwidth is limited, reuse is high, and stragglers matter.
The architecture view shows why this is more than an offload manager. Each shard restores local KV, exchanges lightweight boundary state, and pipelines lower layers upward instead of waiting for the full restore to finish.
Key papers behind the scheduling, offload, and disaggregation story.