AI Post Transformers · Episode Companion

Second-Order Optimization Meets Runtime Scheduling at Scale

Asteria is a runtime, not a new optimizer: it leaves Shampoo/SOAP math untouched and instead rebuilds where curvature state lives, when the cubic-cost linear algebra runs, and how nodes stay in sync — attacking three "physical walls" that have kept curvature-aware training off the LLM critical path.
arXiv:2605.16184 Oxford · May 2026 Lu, Zhang, Yang, Armour Original interactive viz ↗

Lineage: from K-FAC to a runtime, not a new optimizer

Click a node for the one-line pitch. The math barely changes after 2018 — what changes is who's willing to engineer around it.

Three physical walls (why Shampoo never left the paper)

Fill level = how binding each wall is at LLM scale, per the paper's framing. Hover a wall for the mechanism.

Capacity Overlap Consensus

Where second-order state physically lives

Toggle between a generic ZeRO-Offload-style policy and Asteria's asymmetric tiering, which matches placement to how each tensor is produced and consumed.

Hook-orchestrated shadow-state pipeline

The O(d³) inverse-root runs on CPU workers and stages results on a low-priority shadow CUDA stream — it only fills GPU slack the scheduler already has, never competing with the primary stream.

Step time: DGX Spark, single node

Toggle between the raw spike (every 10 steps, native methods' synchronous eigendecomposition) and Asteria's steady-state overhead over AdamW.

SoC energy (% of AdamW)

Single DGX Spark unit, 50 steps, one seed — treat as directional, not conclusive (see episode debate on statistical power).

Staleness budget S sweep

Runtime gain plateaus by S=3–5; final eval loss barely moves across S=1..10. Oxford settles on S=5, then reuses it at 1B/7B without re-sweeping.

The Kronecker-factorization trick

Toggle to see why nobody stores the true Hessian: it's replaced with two much smaller per-dimension factors L and R.

Optimizer state memory: AdamW vs Shampoo-family (illustrative)

AdamW's state scales ~linearly with parameter count N. Shampoo/SOAP-style preconditioners scale far worse per tensor — the gap this paper is built to route around, not close.

References