AI Post Transformers Interactive Visualization arXiv:2311.18677

Splitwise: Phase-Split LLM Inference

A visual companion focused on one systems claim: prefill and decode are different enough that forcing them onto the same hardware can waste throughput, power, and money. The page draws that asymmetry directly as flows, heatmaps, and fleet-level tradeoff charts.
Primary idea
Phase splitting
Prefill
Compute-heavy
Decode
KV + bandwidth
Operator lens
Power-normalized
Paper
Splitwise
University of Washington + Microsoft. Generative inference is split at the prompt-to-first-token boundary.
Paper family
ORCA / PagedAttention / SARATHI / DistServe
Scheduling, KV-cache layout, batch shaping, and disaggregation all target the same serving asymmetry from different angles.
Known arXiv IDs
2311.18677
2401.09670
2602.21548
2603.13358
Additional transcript scan for DDDD.DDDDD found no extra arXiv IDs beyond those listed above.
Headline framing
Different GPU for different phase
Fast FLOP-rich devices for prefill bursts; cheaper or lower-power devices for decode when memory dominates.

Inference Split Map

The serving path is drawn as a baton pass. Prompt tokens create the KV suitcase on one machine; decode resumes on a different machine tuned for repeated cache reads.

Prefill flow Decode flow Transfer state
Hover nodes and links to see phase-specific pressure: parallel matrix compute in prefill, then sequential token generation dominated by KV traffic.

Why a split exists at all

One stage saturates compute; the other keeps revisiting memory. That is the asymmetry the paper treats as operational, not cosmetic.

This radial profile uses illustrative values to show why one hardware profile does not necessarily serve both phases well.

Compute vs Memory Heatmap

Rows show serving situations. Columns show where pressure lands. Hover cells to inspect how prompt length, output length, and reuse assumptions bend the balance.

Cool Warm Hot
Mock values summarize the episode’s framing: long prompts push prefill harder; long generations and weak reuse keep decode memory-hot.

Step-through decode mechanics

Move one token at a time. Each step adds little fresh compute but revisits a growing cache. The phase stays serial even when batching is clever.

Interactive element: step buttons progressively reveal KV growth, bandwidth pressure, and first-token vs nth-token behavior.

Provisioning Toggle

Switch between equal-box and equal-power views. The visual point is not one universal winner, but how the answer changes once operators optimize under power or cost budgets.

Illustrative numbers echo the episode: homogeneous fleets gain from role isolation; heterogeneous fleets gain more when hardware is matched to phase bottlenecks.

Second-token penalty vs network quality

The handoff is only cheap if the back-plane is strong. This chart visualizes how quickly the story deteriorates as transfer latency and congestion rise.

The gap between “paper-friendly” interconnect and “messy shared fabric” is one of the main limits emphasized in the discussion.

Serving Systems Position Map

These papers attack the same asymmetry from different levels: scheduler, memory system, transfer path, or hardware placement.

Click a node to highlight its neighborhood. Splitwise sits farthest toward cluster-level disaggregation; ORCA, PagedAttention, and SARATHI move earlier along the stack.

When phase splitting looks compelling

These mini-panels compress the episode’s caution: phase splitting is strongest when prompt reuse is weak, fabrics are fast, and operators can choose hardware pools deliberately.

This final panel is a deployment guide, not a theorem. Heavy reuse, weak networking, or already-underutilized fleets shrink the payoff.

References

Compact paper trail for the systems cluster around prefill, decode, KV movement, and disaggregation.

Splitwise: Efficient generative LLM inference using phase splitting
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
Not All Prefills Are Equal: PPD Disaggregation for Multi-turn LLM Serving
ORCA, PagedAttention, and SARATHI
Key serving antecedents discussed in the episode for iteration-level scheduling, KV layout, and chunked prefills.
Mooncake, KVLink, Arrow, WindServe, SpecExec, staged speculative decoding
Later work stressing transfer paths, reuse, adaptivity, and the possibility that decode itself becomes more compute-like over time.