Topology: One Saturated Path, One Idle Path
Baseline PD-disaggregation puts all persistent KV loads onto prefill-side storage NICs. DualPath activates decode-side storage bandwidth, then forwards KV data over RDMA to prefill engines.
Short-Append Heatmap
Each row is an agentic session. Early context persists while tiny turn-by-turn appends accumulate, which drives cache reuse higher over time.
Session Growth Profile
Context depth rises steadily while per-turn additions remain small. This is the shape that converts inference into a storage retrieval problem.
Four-Signal Scheduler
The routing decision is not static. DualPath reacts to prefill GPU load, decode GPU load, prefill storage NIC pressure, and decode-side NIC headroom.
Step-by-Step Request Routing
DualPath adds only one structural idea: let decode nodes pull KV from storage first, then forward it over RDMA when that relieves the true bottleneck.
Throughput Comparison
Mock bars track the paper’s reported magnitude: roughly 1.87× offline and 1.96× online average throughput improvement across three unnamed models.
NIC Utilization Before / After
The main effect is not higher GPU math occupancy by itself. It is moving the hot spot away from a single prefill storage interface.
References
Compact pointers to the papers and prior podcast episodes framing the storage-heavy systems context.
arXiv:2602.21548
Scholar search
Scholar search
Scholar search
Scholar search
Scholar search
Scholar search
podcast.do-not-panic.com