The serving path is drawn as a baton pass. Prompt tokens create the KV suitcase on one machine; decode resumes on a different machine tuned for repeated cache reads.
One stage saturates compute; the other keeps revisiting memory. That is the asymmetry the paper treats as operational, not cosmetic.
Rows show serving situations. Columns show where pressure lands. Hover cells to inspect how prompt length, output length, and reuse assumptions bend the balance.
Move one token at a time. Each step adds little fresh compute but revisits a growing cache. The phase stays serial even when batching is clever.
Switch between equal-box and equal-power views. The visual point is not one universal winner, but how the answer changes once operators optimize under power or cost budgets.
The handoff is only cheap if the back-plane is strong. This chart visualizes how quickly the story deteriorates as transfer latency and congestion rise.
These papers attack the same asymmetry from different levels: scheduler, memory system, transfer path, or hardware placement.
These mini-panels compress the episode’s caution: phase splitting is strongest when prompt reuse is weak, fabrics are fast, and operators can choose hardware pools deliberately.
Compact paper trail for the systems cluster around prefill, decode, KV movement, and disaggregation.