← All episodes On-Policy Self-Distillation for Advanced LLM Reasoning

On-Policy Self-Distillation for Advanced LLM Reasoning

Feb 6, 2026
On-policy distillation improves LLM reasoning by using a teacher model to provide dense, token-level feedback on the student's own samples. Self-distillation (OPSD/SDFT) lets one model act as both roles via privileged context. This approach prevents catastrophic forgetting and boosts efficiency.