HyperOffload moves the offload/prefetch decision out of a reactive runtime and into the compiler: the computation graph gets scheduling nodes baked in ahead of time, targeting terabyte-scale shared-memory "SuperNode" hardware.
Reactive Runtime vs. Graph-Driven Scheduling
Top chain: today's reactive runtimes only see memory pressure after it happens, so prefetch decisions arrive late. Bottom chain: HyperOffload inserts offload/prefetch as first-class graph nodes at compile time, before execution ever starts.
SuperNode Hardware Target
Each NPU shares a terabyte-scale remote memory pool. HyperOffload's scheduler decides which tensors move across that link and when — the same link a reactive runtime only reacts to after a stall.
The paper's headline number — 26% peak memory reduction — comes from exactly one config: DeepSeek-V3 with NSA, entire KV cache pushed to remote memory. Table 3's own text says the reduction "closely matches the KV cache size itself."
Does the Headline Number Need a Scheduler?
KV cache share of total memory footprint vs. reported peak memory reduction — they track almost 1:1, which is what you'd expect from moving the whole cache off-device, scheduler or not.
Same Headline, Different Amount of Real Work
Judged on peak memory reduction alone, a naive always-offload policy looks nearly identical to HyperOffload. Click "Full Picture" to see where the compiler-driven schedule actually separates itself.
This is where the scheduler earns its keep: fragmentation stalls disappear entirely, and throughput holds up as D2H bandwidth gets squeezed — neither of those follows from "just offload the KV cache."
Bandwidth-Robustness Curve
HyperOffload's gain over the non-offloading baseline grows as D2H bandwidth tightens — the scheduler is doing more useful work exactly when hardware is more constrained.
Stall Pattern: Reactive vs. Scheduled
Each cell is one execution timestep. Hot cells mark a memory-defragmentation stall. The reactive runtime accumulates them across the run; the compile-time schedule avoids them by construction. Hover a cell for detail.
Section 3.1 opens the paper with a motivating anecdote: reactive prefetching on LLaMA3-8B / Ascend 910C balloons from 5.5s to 15s — a 2.7× slowdown. Section 7 never re-runs that exact scenario through HyperOffload.
Three Conditions, Two Ever Compared
Every Section 7 result compares HyperOffload against a non-offloading baseline. The reactive, runtime-driven prefetching condition that opens the paper is never put side-by-side with HyperOffload directly.
Does Figure 6 Settle It?
Citation Gap: Closest Prior Art, Missing
| Related work | Cited? | Engaged as a competing approach? |
|---|---|---|
| ZeRO | Cited | — |
| ZeRO-Offload | Cited | — |
| ZeRO-Infinity (2021) — tiered NVMe/host offload, closest prior art | Missing | N/A — absent entirely |
| PagedAttention (Kwon et al., SOSP 2023) — block-based virtual KV paging | Cited once | Not engaged — only used to justify "KV caches dominate memory" |
Every result in the paper runs on Ascend NPUs through MindSpore — SJTU and Huawei's own stack, end to end. There's no evidence yet that graph-level offload scheduling ports to CUDA graphs or PyTorch.
Validated Stack vs. Untested Stack
Solid, lit boxes: what the paper actually measures. Dashed, dimmed boxes: the dominant industry stack this has never been tried on.