Click a node for the one-line pitch. The math barely changes after 2018 — what changes is who's willing to engineer around it.
Fill level = how binding each wall is at LLM scale, per the paper's framing. Hover a wall for the mechanism.
Toggle between a generic ZeRO-Offload-style policy and Asteria's asymmetric tiering, which matches placement to how each tensor is produced and consumed.
The O(d³) inverse-root runs on CPU workers and stages results on a low-priority shadow CUDA stream — it only fills GPU slack the scheduler already has, never competing with the primary stream.
Toggle between the raw spike (every 10 steps, native methods' synchronous eigendecomposition) and Asteria's steady-state overhead over AdamW.
Single DGX Spark unit, 50 steps, one seed — treat as directional, not conclusive (see episode debate on statistical power).
Runtime gain plateaus by S=3–5; final eval loss barely moves across S=1..10. Oxford settles on S=5, then reuses it at 1B/7B without re-sweeping.
Toggle to see why nobody stores the true Hessian: it's replaced with two much smaller per-dimension factors L and R.
AdamW's state scales ~linearly with parameter count N. Shampoo/SOAP-style preconditioners scale far worse per tensor — the gap this paper is built to route around, not close.