AI Post Transformers · Episode Companion

StrataCL: Fabric-Native Zero-Copy Comms for AI Supernodes

arXiv:2607.26444 Peking University · ICT-CAS · UCAS · SJTU · Huawei Testbed: CloudMatrix384 1.4x iteration speedup · 512 dies

StrataCL gives collectives and MoE routing true zero-copy access to application buffers on unified-memory fabrics, without breaking PyTorch / SGLang compatibility. The core trick — registration-on-allocation — exploits the multi-second gap between physical memory allocation and first use to move buffer registration off the critical path entirely.

1.4x
iteration speedup, 512 dies
30-45%
comm overhead, training
50%+
overhead at scale
2.6s
alloc-to-first-use gap

Communication overhead by regime

Share of end-to-end time spent moving bytes, not computing on them

What "unified memory fabric" means

CloudMatrix384: 384 Ascend NPUs / 768 dies over the Unified Bus
Scale-up fabric (Unified Bus), not scale-out RDMA (InfiniBand). Registration collapses from a multi-ms RDMA handshake to a microsecond address-mapping update.

Registration timing: the exploit

Toggle to compare when registration happens relative to first communication use
Framework allocators (KV-cache, graph capture, warm-up) leave a gap of at least 2.6 seconds between physical allocation and first touch by a comm operator. StrataCL intercepts the allocation call and registers asynchronously in the background — by the time an operator needs the buffer, registration is already done.

Shadow virtual addressing

Every NPU gets a disjoint VA range at bootstrap; a buffer's address is mirrored identically across all peers
NPU 0 allocates at virtual address 0x7f2a…. StrataCL mirrors that exact VA on every peer, mapped through the Unified Bus back to NPU 0's physical HBM. A peer issues a direct load/store at the identical address — no per-buffer lookup table, no translation round trip.

Registration latency, by path

Log-scale comparison — lower is better

VMM remap handling

PyTorch expandable-segment allocator maps new physical pages onto a reserved VA range on demand
<4%
MoE batches trigger remap
<0.6%
net overhead vs full VA

Ablation: where the speedup comes from

Toggle between inference throughput stack-up and training / operator-level gains

Workload-balanced core partitioning

Fastest-to-slowest core finish-time gap, naive vs balanced

SDMA offloading trade

A little latency, a lot of freed compute cores

Where the story gets shakier: scale & baseline

Bandwidth edge over the HCCL-zerocopy baseline as rank count climbs
StrataCL's own ceiling is 768 dies; evaluation stops at 512 (training) and 256 (microbenchmarks). Within that range it already loses to HCCL-zerocopy by ~6% at large payloads and loses to ring past 16 MiB — full-mesh isn't free at scale.

Full-mesh vs. ring decomposition

Toggle communication pattern — ring minimizes concurrent peers, full-mesh fires every slice in one logical step
data transfer synchronization wait

Fabric topology tiers

Die-to-die is ~10x faster than a cross-node hop — hover a cell

MoE dispatch / combine

Token routing changes every batch — buffer sizes are never stable

Ring synchronization tax

At 64 KiB payloads, ring spends the majority of its time waiting, not moving bytes

References

  1. Hu, Qin, Wang, Liu, Li, Wang, Hu, Hu, Wang, Li, Zhou, Bao, Sun, Zhao, Cui, Xie, Wang — StrataCL: Fabric-Native Communication Library for Production Supernodes, 2026. arXiv:2607.26444
  2. von Eicken, Basu, Buch, Vogels — U-Net: A User-Level Network Interface for Parallel and Distributed Computing, 1995. Scholar
  3. Kalia, Kaminsky, Andersen — Design Guidelines for High Performance RDMA Systems, 2016. Scholar
  4. Wang, Venkataraman, Phanishayee, Thelin, Devanur, Stoica — Blink: Fast and Generic Collectives for Distributed ML, 2020. Scholar
  5. NVIDIA — NVSHMEM: GPU-initiated, PGAS-style one-sided communication library, 2016–. Scholar
  6. Si, Balaji, Chen, et al. — Collective Communication for 100k+ GPUs, 2025. Scholar
  7. Li, Liu, Huang, et al. — SwiftEP: Accelerating MoE Inference with Buffer Fusion and TMA Offloading, NSDI 2026. Scholar
  8. PyTorch Team / NVIDIA — PyTorch Symmetric Memory / NVSHMEM-style same-VA mirrored buffers, 2024–2025. Scholar
  9. Kwon, Li, Zhuang, et al. — Efficient Memory Management for Large Language Model Serving with PagedAttention, SOSP 2023. Scholar