From Static Imitation to Precomputed Judgment
Eleven years of distillation and reasoning-RL lineage, compressed to one line
Hover any node for the full citation. Lightning OPD is the largest node — it doesn't introduce on-policy distillation, it removes the live teacher that every prior version of it required.
Where This Sits in Post-Training
Pretrain → SFT → the second-stage fork this paper lives inside
RLVR checks a final answer and back-propagates one scalar through the whole trace. On-policy distillation scores every token the student produces against a teacher's distribution — Lightning OPD is about how cheaply that scoring can happen.
Signal Density: One Bit vs. One Distribution Per Token
Why on-policy distillation is pitched as gentler than RLVR
Each cell is one token in a reasoning trace, colored by how much training signal it carries. RLVR lights up only the last token — the verified answer. OPD scores every token against the teacher's distribution.
Standard OPD vs. Lightning OPD
The infrastructure difference the whole paper is about
The Lightning OPD Recipe, Stage by Stage
Click a stage to see what happens in it
The Drift Bound the Whole Method Leans On
Theorem 3.5: online/offline gradient gap ∝ √(χ² divergence between student and π-ref)
Small drift → small bound → frozen rollouts stay valid. The paper validates this for exactly 150 steps (Section 4.1, Fig. 3b). Frontier RLVR/OPD runs go for thousands — whether the bound still holds that far out is untested, not disproven.
Benchmark Parity — and a Few Wins
AIME 2024/2025, HMMT 2025, LiveCodeBench v5/v6
Lightning OPD tracks standard OPD closely and edges past ExOPD across the board — 68.1 vs. ExOPD's 61.0 on AIME 2024 at 4B, for instance.
The Efficiency Headline
GPU-hours to reach these numbers
The 8B run hits 69.9% on AIME 2024 in just 30 GPU-hours — a quarter of what a live-teacher run costs.
Where Offline Stops Being Optional
Qwen3-30B-A3B (MoE, 3B active) on a single 8×H100 node
Standard OPD needs both a 30B teacher and 30B student co-resident for scoring — it OOMs on one node. Lightning OPD's teacher was already consulted and discarded, so it trains that same model to 71.0% on AIME 2024 and 60.8% on LiveCodeBench v5 on hardware a single academic lab could own.
The Teacher-Consistency Ablation
Table 4 — accuracy drop at 8B when SFT teacher ≠ OPD teacher
A mismatched teacher hurts Lightning OPD more than standard OPD — the frozen rollout set does double duty as both reference distribution and sole training data, so a corrupted π-ref poisons both at once.