1.4x
iteration speedup, 512 dies
30-45%
comm overhead, training
50%+
overhead at scale
2.6s
alloc-to-first-use gap
Communication overhead by regime
Share of end-to-end time spent moving bytes, not computing on them
What "unified memory fabric" means
CloudMatrix384: 384 Ascend NPUs / 768 dies over the Unified Bus
Scale-up fabric (Unified Bus), not scale-out RDMA
(InfiniBand). Registration collapses from a multi-ms RDMA
handshake to a microsecond address-mapping update.
Registration timing: the exploit
Toggle to compare when registration happens relative to first communication use
Framework allocators (KV-cache, graph capture, warm-up) leave a
gap of at least 2.6 seconds between physical allocation
and first touch by a comm operator. StrataCL intercepts the
allocation call and registers asynchronously in the background —
by the time an operator needs the buffer, registration is
already done.
Shadow virtual addressing
Every NPU gets a disjoint VA range at bootstrap; a buffer's address is mirrored identically across all peers
NPU 0 allocates at virtual address
0x7f2a…. StrataCL
mirrors that exact VA on every peer, mapped through the Unified
Bus back to NPU 0's physical HBM. A peer issues a direct
load/store at the identical address — no per-buffer lookup
table, no translation round trip.
Registration latency, by path
Log-scale comparison — lower is better
VMM remap handling
PyTorch expandable-segment allocator maps new physical pages onto a reserved VA range on demand
<4%
MoE batches trigger remap
<0.6%
net overhead vs full VA
Ablation: where the speedup comes from
Toggle between inference throughput stack-up and training / operator-level gains
Workload-balanced core partitioning
Fastest-to-slowest core finish-time gap, naive vs balanced
SDMA offloading trade
A little latency, a lot of freed compute cores
Where the story gets shakier: scale & baseline
Bandwidth edge over the HCCL-zerocopy baseline as rank count climbs
StrataCL's own ceiling is 768 dies; evaluation stops at 512
(training) and 256 (microbenchmarks). Within that range it
already loses to HCCL-zerocopy by ~6% at large payloads
and loses to ring past 16 MiB — full-mesh isn't free
at scale.
Full-mesh vs. ring decomposition
Toggle communication pattern — ring minimizes concurrent peers, full-mesh fires every slice in one logical step
data transfer
synchronization wait
Fabric topology tiers
Die-to-die is ~10x faster than a cross-node hop — hover a cell
MoE dispatch / combine
Token routing changes every batch — buffer sizes are never stable
Ring synchronization tax
At 64 KiB payloads, ring spends the majority of its time waiting, not moving bytes