← All episodes Decoupling KL Direction from Rollout Source in LLM Distillation

Decoupling KL Direction from Rollout Source in LLM Distillation

Aug 28, 2026
This episode examines "Decoupling KL and Trajectories," which challenges the assumption that off-policy training must pair with forward KL divergence and on-policy training with reverse KL — a coupling used by DeepSeek-R1, Qwen3, MiMo, and GLM-5 without ever being tested. The hosts unpack the two independent design axes at play: prefix source (teacher-generated versus student-generated rollouts) and KL direction (forward's distribution-covering behavior versus reverse's mode-seeking concentration), showing that standard SFT is actually just forward-KL distillation on teacher text in disguise. The paper's real contribution is exploring two previously unstudied combinations — teacher-prefix reverse KL and student-prefix forward KL — turning an assumed two-option choice into a full four-quadrant design space. Grounding the discussion is a concrete experimental setup using Qwen3-0.6B-Base as a student distilling from Qwen3-4B and Qwen3-8B teachers. Listeners interested in LLM distillation, on-policy training, or exposure bias will find this a sharp dismantling of an industry-wide convention nobody had thought to question.
Sources:
1. Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation — Anhao Zhao, Haoran Xin, Yingqi Fan, Junlong Tong, Wenjie Li, Xiaoyu Shen, 2026
http://arxiv.org/abs/2605.16826
2. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011
https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning
3. Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks — Samy Bengio, Oriol Vinyals, Navdeep Jaitly, Noam Shazeer, 2015
https://scholar.google.com/scholar?q=Scheduled+Sampling+for+Sequence+Prediction+with+Recurrent+Neural+Networks
4. Sequence Level Training with Recurrent Neural Networks — Marc'Aurelio Ranzato, Sumit Chopra, Michael Auli, Wojciech Zaremba, 2016
https://scholar.google.com/scholar?q=Sequence+Level+Training+with+Recurrent+Neural+Networks
5. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes — Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, Olivier Bachem (Google DeepMind), 2024
https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes
6. MiniLLM: Knowledge Distillation of Large Language Models — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2024
https://scholar.google.com/scholar?q=MiniLLM%3A+Knowledge+Distillation+of+Large+Language+Models
7. Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models — Feng Luo, Yu-Neng Chuang, Guanchu Wang, Zicheng Xu, Xiaotian Han, Tianyi Zhang, Vladimir Braverman, 2026
https://scholar.google.com/scholar?q=Demystifying+OPD%3A+Length+Inflation+and+Stabilization+Strategies+for+Large+Language+Models
8. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Stéphane Ross, Geoffrey J. Gordon, J. Andrew Bagnell, 2011
https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning+%28DAgger%29
9. Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting — Xinyuan Chen, Noam Razin, Karthik Narasimhan, Danqi Chen, 2025
https://scholar.google.com/scholar?q=Retaining+by+Doing%3A+The+Role+of+On-Policy+Data+in+Mitigating+Forgetting
Interactive Visualization: Decoupling KL Direction from Rollout Source in LLM Distillation