MegaTrain: Training 120B-Parameter Models on a Single GPU

Zhengqing Yuan, Hanchi Sun, Lichao Sun, Yanfang Ye — Notre Dame & Lehigh University, 2026
arXiv:2604.05091 H200 · 1.5TB host RAM · single GPU 120B params · full precision Full interactive viz →

The four-tier memory hierarchy under a GPU

Bandwidth drops by orders of magnitude at every tier down; capacity grows the same way. PCIe Gen5 (128 GB/s) is the toll booth between HBM and host DDR5 — the dashed line below.

12 bytes per parameter: the Adam tax

BF16 weights (2B) + BF16 gradients (2B) + FP32 first/second moment (4B+4B) = 12 bytes/param, before a single activation is generated. Hover a segment for the breakdown.

Who owns the data: GPU or host?

ZeRO-Offload / ZeRO-Infinity treat host memory and NVMe as a spill buffer for whatever HBM can't hold. MegaTrain inverts this: host RAM is the authoritative home for every parameter and optimizer state; GPU HBM becomes a transient scratchpad for one layer at a time.

Three CUDA streams, ping-ponging across PCIe

While the compute stream multiplies layer N, H2D is already prefetching layer N+1's weights and D2H is draining layer N−1's gradients back to host RAM. Step through it, or hit play.

Layer 1 / 12

Sustained TFLOPS from 7B to 120B

Triple-digit throughput holds as parameter count climbs; host memory footprint scales roughly linearly instead of exploding.

Head-to-head: who else claimed this

What was actually verified, by scale

Correctness (loss curves, downstream accuracy vs full-precision baseline) exists only at 7B/14B on MetaMathQA. Everything from 32B up is TFLOPS and memory footprint only. Hover a cell.

Prior art timeline: inversion, streaming, and offloading

The GPU-as-transient-cache idea and layer-streaming mechanics both predate this paper; what's argued as new is bundling CPU-resident full Adam + stateless templates into one working system at this scale.

References