This episode examines why on-policy distillation (OPD) of large language models can fail catastrophically even when the teacher model is objectively stronger by every benchmark — a 7-billion-parameter teacher completely failed to improve a 1.5-billion-parameter student, while a smaller, weaker teacher succeeded. The discussion traces OPD's mechanics: unlike classic distillation, which trains students on teacher-generated text and suffers from exposure bias, OPD has the student generate its own rollouts and scores them against the teacher's token-level probability distribution via reverse KL divergence, yielding a dense per-token reward without a verifier. The central finding is that teacher-student "overlap ratio" — how much the teacher's likely next tokens actually match the student's own candidate set — determines whether that dense signal teaches anything, meaning benchmark strength and teachability are fundamentally different axes. The paper's three-part structure (phenomenology, mechanism, recipe) is illustrated through a controlled comparison of two similarly-scored Qwen3-4B teacher variants, isolating thinking-pattern compatibility as the real driver of distillation success. Listeners working on post-training or distillation pipelines will find this a direct challenge to the common assumption that upgrading the teacher model is a safe, unconditional improvement.
Sources:
1. Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe — Yaxuan Li, Yuxin Zuo, Bingxiang He, Jinqian Zhang, Chaojun Xiao, Cheng Qian, Tianyu Yu, Huan-ang Gao, Wenkai Yang, Zhiyuan Liu, Ning Ding, 2026
http://arxiv.org/abs/2604.130162. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning — Stéphane Ross, Geoffrey Gordon, J. Andrew Bagnell, 2011
https://scholar.google.com/scholar?q=A+Reduction+of+Imitation+Learning+and+Structured+Prediction+to+No-Regret+Online+Learning3. MiniLLM: Knowledge Distillation of Large Language Models — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2023
https://scholar.google.com/scholar?q=MiniLLM%3A+Knowledge+Distillation+of+Large+Language+Models4. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes (GKD) — Rishabh Agarwal, Nino Vieillard, and colleagues (Google DeepMind), 2024
https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes+%28GKD%295. Qwen3 Technical Report — An Yang and the Qwen Team (Alibaba), 2025
https://scholar.google.com/scholar?q=Qwen3+Technical+Report6. On-policy distillation of language models: Learning from self-generated mistakes (MiniLLM) — Yuxian Gu, Li Dong, Furu Wei, Minlie Huang, 2023
https://scholar.google.com/scholar?q=On-policy+distillation+of+language+models%3A+Learning+from+self-generated+mistakes+%28MiniLLM%297. Distillation scaling laws — Dan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram, Etai Littwin, Russ Webb, 2025
https://scholar.google.com/scholar?q=Distillation+scaling+laws8. On the efficacy of knowledge distillation — Jang Hyun Cho, Bharath Hariharan, 2019
https://scholar.google.com/scholar?q=On+the+efficacy+of+knowledge+distillation9. Small models struggle to learn from strong reasoners — Yuetai Li, Xiang Yue, Zhangchen Xu, Fengqing Jiang, Luyao Niu, Bill Yuchen Lin, Bhaskar Ramasubramanian, Radha Poovendran, 2025
https://scholar.google.com/scholar?q=Small+models+struggle+to+learn+from+strong+reasoners10. On-policy distillation (Thinking Machines Lab blog) — Kevin Lu and Thinking Machines Lab, 2025
https://scholar.google.com/scholar?q=On-policy+distillation+%28Thinking+Machines+Lab+blog%2911. Learning beyond teacher: Generalized on-policy distillation with reward extrapolation — Wenkai Yang, Weijie Liu, Ruobing Xie, Kai Yang, Saiyong Yang, Yankai Lin, 2026
https://scholar.google.com/scholar?q=Learning+beyond+teacher%3A+Generalized+on-policy+distillation+with+reward+extrapolationInteractive Visualization: On-Policy Distillation: Why a Stronger Teacher Can Backfire