AI Post Transformers · Interactive Visualization

CacheFlow and 3D-Parallel KV Cache Restoration

A visual companion about long-context LLM serving where the bottleneck shifts from token generation to recovering old KV state fast enough to cut time-to-first-token.

arXiv 2604.25080 Transcript IDs none beyond 2604.25080 Focus TTFT, scheduling, restore overlap Viz URL live companion
Headline gain
10–62%
TTFT improvement reported across workloads
Optimization axis
3D
Tokens · Layers · GPUs
Core scheduler
2-pointer
Shift the cut between recompute and transfer
Real bottleneck
TTFT
First token beats raw decode speed

Restore space is a latency map, not a binary choice

Switch the mode to see how the same request hits different bottlenecks. The diagram emphasizes where waiting time accumulates before the first token appears.

compute path transfer path stalled boundary TTFT pain point

What the visual is saying

Mock trace: longer prefixes push pure recompute up sharply, while pure I/O gets noisy under contention. Hybrid wins by flattening the worst tail.

Token-layer heatmap and the moving cut frontier

Late tokens cost more to recompute because they drag a larger history window. Hover the matrix, then step the scheduler to watch the recompute/load boundary move under contention.

cold / cheap medium hot / expensive

Batch-level split

Reported win range only appears when restores are heterogeneous

Use the hardware and traffic toggles. This chart is mock data shaped to match the episode’s claim: the scheduler looks best when bandwidth is limited, reuse is high, and stragglers matter.

Lower is better. Bars show first-token latency; the overlay line shows relative speedup versus pure I/O restore.

Latency surface

This surface highlights where control logic pays for itself: large prompts plus batch contention produce the steepest penalty for naive policies.

3D parallelism across tokens, layers, and GPU shards

The architecture view shows why this is more than an offload manager. Each shard restores local KV, exchanges lightweight boundary state, and pipelines lower layers upward instead of waiting for the full restore to finish.

local compute KV movement boundary hidden state staged layer pipeline

References and nearest context

Key papers behind the scheduling, offload, and disaggregation story.