AI Post Transformers · Episode Companion

Distributed Weight Data Parallelism Cuts LLM Inference Stalls

DWDP keeps GPUs fully data-parallel while each one asynchronously prefetches missing MoE expert weights from peers over NVLink copy engines — timed to hide behind ongoing compute, instead of forcing every rank through a synchronized layer-boundary barrier.

GPU Rank Architecture — DEP Baseline vs DWDP

Four GPU ranks in an NVL72 group. Attention weights are always replicated per-rank (data-parallel). Toggle how the MoE expert layer moves weights across ranks.

Why this matters

Every existing model-parallel strategy — expert, tensor, pipeline — forces GPUs to hit a synchronization barrier at each layer boundary. Under ordinary workload imbalance, the paper's own baseline burns roughly 12% of total inference time waiting there.

Smarter scheduling (cache-aware, load-aware routing) shrinks the imbalance feeding the wait, but the barrier itself — the wait — stays architecturally required.

The DWDP idea in one line

Stay data-parallel. Instead of a synchronized collective, each GPU pulls whichever expert weights it's missing directly from a peer's memory via cudaMemcpyAsync over NVLink — on the copy engine, never touching the SMs, timed to finish before it's needed two blocks later.

Compute vs. Synchronization Wait

Coefficient of variation (CV) in sequence length across ranks drives idle time. Toggle imbalance level.

Per-Rank / Per-Layer Wait Heatmap

Synchronization wait time as a fraction of layer latency, 4 ranks × 6 layers. Hover a cell for its value.

low wait moderate high wait

Two sources of straggler pressure

Request-level imbalance — ranks get different sequence lengths and KV-cache hit rates. Weight-level imbalance — MoE routing skews so some ranks serve hotter experts. Both feed the same synchronized all-to-all barrier in DEP.

Double-Buffered Prefetch Timeline

While a rank computes layer l's MoE block and layer l+1's attention block, the copy engine prefetches layer l+1's remote experts into the alternate buffer. Two full compute blocks hide one fetch.

SM compute (attn / MoE) copy-engine prefetch (buffer B) idle

Roofline Crossover

Per-layer latency: DWDP = max(compute, prefetch); DEP = compute + all-to-all. DWDP overtakes DEP around 16K input tokens at batch size 1.

Expert Buffer Layout

Local and remote expert weights must reach the groupedGEMM kernel as one contiguous buffer. Toggle the fix.

Source-Rank Contention

Multiple destinations can converge on one source rank mid-pull. In a 4-rank group there's an ~11% chance three requests stack on a single source at once — fixed with time-division multiplexed slicing across destinations.

Speedup vs. Input Length

TPS/GPU stays a steady ~1.09–1.11× across input lengths. TTFT speedup is largest at short sequences (small compute window still hides most of the fetch) and dips as sequences grow, before climbing again with a bigger token budget.

TPS per GPU speedup TTFT speedup

Gross → Net Latency Reduction

Removing sync yields a 21.86% gross reduction in iteration latency. Buffer-merge overhead and attention-block slowdown eat into it, leaving 11.69% realized net.

DWDP3 vs DWDP4

DWDP3 tracks DWDP4 almost exactly on TPS/GPU — but regresses on TTFT (0.86×, below the 1.0 baseline), while DWDP4 stays a net win on both.

Serving-System TTFT Trade-off

8.8% average TPS/GPU gain comes from needing fewer context GPUs — which lowers the context stage's service rate and queues requests, inflating TTFT. Toggle the TPS-per-user regime.

References