Bandwidth drops by orders of magnitude at every tier down; capacity grows the same way. PCIe Gen5 (128 GB/s) is the toll booth between HBM and host DDR5 — the dashed line below.
BF16 weights (2B) + BF16 gradients (2B) + FP32 first/second moment (4B+4B) = 12 bytes/param, before a single activation is generated. Hover a segment for the breakdown.
ZeRO-Offload / ZeRO-Infinity treat host memory and NVMe as a spill buffer for whatever HBM can't hold. MegaTrain inverts this: host RAM is the authoritative home for every parameter and optimizer state; GPU HBM becomes a transient scratchpad for one layer at a time.
While the compute stream multiplies layer N, H2D is already prefetching layer N+1's weights and D2H is draining layer N−1's gradients back to host RAM. Step through it, or hit play.
Triple-digit throughput holds as parameter count climbs; host memory footprint scales roughly linearly instead of exploding.
Correctness (loss curves, downstream accuracy vs full-precision baseline) exists only at 7B/14B on MetaMathQA. Everything from 32B up is TFLOPS and memory footprint only. Hover a cell.
The GPU-as-transient-cache idea and layer-streaming mechanics both predate this paper; what's argued as new is bundling CPU-resident full Adam + stateless templates into one working system at this scale.