AI Post Transformers • Visual Companion

DualPath Breaks Storage Bandwidth Bottleneck in Agentic Inference

This page visualizes the paper’s core asymmetry: prefill storage NICs are pinned at 100% while decode-side storage links sit mostly idle. The visuals focus on data movement, cache reuse, scheduler routing, and the throughput gains that come from exploiting hardware already present in PD-disaggregated clusters.

Paper
Yongtong Wu et al. • February 26, 2026
Peking University • Tsinghua University • DeepSeek-AI
Episode Lens
Short-append agentic sessions push KV-cache hit rates above 95%, moving the bottleneck from GPU arithmetic to storage I/O.
Headline Result
Up to 1.87× offline throughput and 1.96× online throughput by adding a decode-side cache loading path and routing dynamically.

Topology: One Saturated Path, One Idle Path

Baseline PD-disaggregation puts all persistent KV loads onto prefill-side storage NICs. DualPath activates decode-side storage bandwidth, then forwards KV data over RDMA to prefill engines.

active transfer decode-assisted route bottleneck / hot link compute RDMA fabric
Prefill Storage NIC 100%
Observed production saturation driving H100s down to roughly 40% compute utilization.
Decode Storage NIC Idle
Equivalent hardware exists on decode nodes but is underused in the standard design.
KV Cache Hit Rate 95%+
Real production agentic traces make cache loading, not cache recompute, the dominant path.
DualPath Objective 2 Paths
A scheduler chooses between direct storage→prefill and storage→decode→prefill.

Short-Append Heatmap

Each row is an agentic session. Early context persists while tiny turn-by-turn appends accumulate, which drives cache reuse higher over time.

Hover cells to inspect turn-level context reuse and append intensity.

Session Growth Profile

Context depth rises steadily while per-turn additions remain small. This is the shape that converts inference into a storage retrieval problem.

Four-Signal Scheduler

The routing decision is not static. DualPath reacts to prefill GPU load, decode GPU load, prefill storage NIC pressure, and decode-side NIC headroom.

Step-by-Step Request Routing

DualPath adds only one structural idea: let decode nodes pull KV from storage first, then forward it over RDMA when that relieves the true bottleneck.

Throughput Comparison

Mock bars track the paper’s reported magnitude: roughly 1.87× offline and 1.96× online average throughput improvement across three unnamed models.

NIC Utilization Before / After

The main effect is not higher GPU math occupancy by itself. It is moving the hot spot away from a single prefill storage interface.

References

Compact pointers to the papers and prior podcast episodes framing the storage-heavy systems context.

DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference Wu et al., 2026
arXiv:2602.21548
PagedAttention Kwon et al., 2023
Scholar search
Mooncake Qin et al., 2024
Scholar search
DistServe Zhong et al., 2024
Scholar search
Splitwise Patel et al., 2024
Scholar search
CacheGen Liu et al., 2024
Scholar search
AgentBench Liu et al., 2023
Scholar search
AI Post Transformers Related Episodes Bidaw • SYMPHONY • CXL-SpecKV • Memory Wall • Tempo
podcast.do-not-panic.com