NVIDIA · 2026 arXiv:2604.13010 Post-Training / Distillation

Lightning OPD: Precomputing the Teacher for Cheaper Reasoning Distillation

On-policy distillation gives denser training signal than RLVR, but has always needed a live teacher model running beside the student. Lightning OPD shows the student's rollout distribution drifts only modestly from its SFT start — so you can capture the teacher's judgment once, offline, and train against the fossil.

Yecheng Wu · Song Han · Hai Cai — NVIDIA, 2026

From Static Imitation to Precomputed Judgment

Eleven years of distillation and reasoning-RL lineage, compressed to one line

Distillation lineage Verification / reward models RLVR lineage On-policy distillation

Hover any node for the full citation. Lightning OPD is the largest node — it doesn't introduce on-policy distillation, it removes the live teacher that every prior version of it required.

Where This Sits in Post-Training

Pretrain → SFT → the second-stage fork this paper lives inside

RLVR checks a final answer and back-propagates one scalar through the whole trace. On-policy distillation scores every token the student produces against a teacher's distribution — Lightning OPD is about how cheaply that scoring can happen.

Signal Density: One Bit vs. One Distribution Per Token

Why on-policy distillation is pitched as gentler than RLVR

Each cell is one token in a reasoning trace, colored by how much training signal it carries. RLVR lights up only the last token — the verified answer. OPD scores every token against the teacher's distribution.

Standard OPD vs. Lightning OPD

The infrastructure difference the whole paper is about

Teacher consistency requirement: the same teacher model must generate the stage-1 SFT trajectories AND the stage-2 reference log-probabilities. Thinking Machines Lab's public pipeline violates this — SFT on QwQ-32B data, distilled against Qwen3-32B — which the paper's Theorems 3.8/3.9 show injects gradient bias into both online and offline OPD.

The Lightning OPD Recipe, Stage by Stage

Click a stage to see what happens in it

The Drift Bound the Whole Method Leans On

Theorem 3.5: online/offline gradient gap ∝ √(χ² divergence between student and π-ref)

Small drift → small bound → frozen rollouts stay valid. The paper validates this for exactly 150 steps (Section 4.1, Fig. 3b). Frontier RLVR/OPD runs go for thousands — whether the bound still holds that far out is untested, not disproven.

Benchmark Parity — and a Few Wins

AIME 2024/2025, HMMT 2025, LiveCodeBench v5/v6

Standard OPD (live teacher) Lightning OPD (offline) ExOPD

Lightning OPD tracks standard OPD closely and edges past ExOPD across the board — 68.1 vs. ExOPD's 61.0 on AIME 2024 at 4B, for instance.

The Efficiency Headline

GPU-hours to reach these numbers

The 8B run hits 69.9% on AIME 2024 in just 30 GPU-hours — a quarter of what a live-teacher run costs.

Where Offline Stops Being Optional

Qwen3-30B-A3B (MoE, 3B active) on a single 8×H100 node

Standard OPD needs both a 30B teacher and 30B student co-resident for scoring — it OOMs on one node. Lightning OPD's teacher was already consulted and discarded, so it trains that same model to 71.0% on AIME 2024 and 60.8% on LiveCodeBench v5 on hardware a single academic lab could own.

The Teacher-Consistency Ablation

Table 4 — accuracy drop at 8B when SFT teacher ≠ OPD teacher

A mismatched teacher hurts Lightning OPD more than standard OPD — the frozen rollout set does double duty as both reference distribution and sole training data, so a corrupted π-ref poisons both at once.

Open question: parity holds at 150 steps, a horizon short enough that "drift stays small" may just mean neither method has moved far from the SFT ceiling yet — not that the offline trick scales to thousand-step runs.

References

  1. Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation — Yecheng Wu, Song Han, Hai Cai (NVIDIA, 2026)
  2. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes — Agarwal, Vieillard, Zhou, et al. (Google DeepMind, ICLR 2024)
  3. On-Policy Distillation — Kevin Lu et al. (Thinking Machines Lab, 2025)
  4. Distilling the Knowledge in a Neural Network — Hinton, Vinyals, Dean (2015)
  5. Tulu 3: Pushing Frontiers in Open Language Model Post-Training — Lambert, Morrison, Pyatkin, et al. (AI2, 2024)
  6. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — DeepSeek-AI (2025)
  7. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models — Shao, Wang, Zhu, et al. (DeepSeek-AI, 2024)
  8. Let's Verify Step by Step — Lightman, Kosaraju, Burda, et al. (OpenAI, 2023)
  9. On-policy distillation (Thinking Machines Lab: Connectionism) — Kevin Lu (2025)
  10. Does RL Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? — Yue, Chen, Lu, et al. (Tsinghua, 2025)
  11. RL's Razor: Why Online Reinforcement Learning Forgets Less — Shenfeld, Pari, Agrawal (MIT, 2025)
  12. Rethinking On-Policy Distillation: Phenomenology, Mechanism, and Recipe — Chen, Guo, Tan, et al. (DeepSeek, 2026)
  13. Learning Beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation (ExOPD) — Yang, Liu, Xie, et al. (2026)