DWDP keeps GPUs fully data-parallel while each one asynchronously prefetches missing MoE expert weights from peers over NVLink copy engines — timed to hide behind ongoing compute, instead of forcing every rank through a synchronized layer-boundary barrier.
Four GPU ranks in an NVL72 group. Attention weights are always replicated per-rank (data-parallel). Toggle how the MoE expert layer moves weights across ranks.
Every existing model-parallel strategy — expert, tensor, pipeline — forces GPUs to hit a synchronization barrier at each layer boundary. Under ordinary workload imbalance, the paper's own baseline burns roughly 12% of total inference time waiting there.
Smarter scheduling (cache-aware, load-aware routing) shrinks the imbalance feeding the wait, but the barrier itself — the wait — stays architecturally required.
Stay data-parallel. Instead of a synchronized collective, each GPU pulls whichever expert weights it's missing directly from a peer's memory via cudaMemcpyAsync over NVLink — on the copy engine, never touching the SMs, timed to finish before it's needed two blocks later.
Coefficient of variation (CV) in sequence length across ranks drives idle time. Toggle imbalance level.
Synchronization wait time as a fraction of layer latency, 4 ranks × 6 layers. Hover a cell for its value.
Request-level imbalance — ranks get different sequence lengths and KV-cache hit rates. Weight-level imbalance — MoE routing skews so some ranks serve hotter experts. Both feed the same synchronized all-to-all barrier in DEP.
While a rank computes layer l's MoE block and layer l+1's attention block, the copy engine prefetches layer l+1's remote experts into the alternate buffer. Two full compute blocks hide one fetch.
Per-layer latency: DWDP = max(compute, prefetch); DEP = compute + all-to-all. DWDP overtakes DEP around 16K input tokens at batch size 1.
Local and remote expert weights must reach the groupedGEMM kernel as one contiguous buffer. Toggle the fix.
Multiple destinations can converge on one source rank mid-pull. In a 4-rank group there's an ~11% chance three requests stack on a single source at once — fixed with time-division multiplexed slicing across destinations.
TPS/GPU stays a steady ~1.09–1.11× across input lengths. TTFT speedup is largest at short sequences (small compute window still hides most of the fetch) and dips as sequences grow, before climbing again with a bigger token budget.
Removing sync yields a 21.86% gross reduction in iteration latency. Buffer-merge overhead and attention-block slowdown eat into it, leaving 11.69% realized net.
DWDP3 tracks DWDP4 almost exactly on TPS/GPU — but regresses on TTFT (0.86×, below the 1.0 baseline), while DWDP4 stays a net win on both.
8.8% average TPS/GPU gain comes from needing fewer context GPUs — which lowers the context stage's service rate and queues requests, inflating TTFT. Toggle the TPS-per-user regime.