arXiv:2605.16826 Qwen3-0.6B student · Qwen3-4B/8B teachers Math reasoning: AIME24 · AMC23 · MATH500 · GSM8K

Decoupling KL Direction from Rollout Source in LLM Distillation

Anhao Zhao, Haoran Xin, Yingqi Fan, Junlong Tong, Wenjie Li, Xiaoyu Shen — Eastern Institute of Technology, Ningbo · HK PolyU · Shanghai Jiao Tong · HKUST Guangzhou. Submitted May 16, 2026. Every production recipe — DeepSeek-R1, Qwen3, MiMo, GLM-5 — locks prefix source to KL direction. This paper unlocks them into a full four-quadrant design space.

The Locked Pair, Unlocked

Two knobs — prefix source (whose text the student trains on) and KL direction (which way the divergence is measured) — are mathematically independent. Every production pipeline still only ever uses two of the four combinations. Click a cell to see what each quadrant actually is.

KL DIRECTION Forward KL Reverse KL PREFIX SOURCE Teacher Student Off-policy SFT = plain distillation (already standard practice) Offline-RL-style distillation objective ★ unstudied before this paper DAgger-style on-policy SFT ★ unstudied before this paper On-Policy Distillation = dense-reward on-policy RL (already standard practice)
Click any quadrant above for the plain-language description.

Pipeline: Where the Two Axes Enter

Teacher Qwen3-4B / 8B Prefix source? teacher text vs student rollout KL direction? forward (cover) vs reverse (seek) 4-cell objective gradient computed Student updated Qwen3-0.6B-Base

The paper's claim: this is two independent forks, not one. Standard SFT is quietly just the top-forward path — forward-KL distillation on teacher-generated text.

Proposition 1 — Same Prefix, Two Different Gradients

Fix a prefix. The gradient of forward KL collapses to plain cross-entropy. The gradient of reverse KL collapses to a REINFORCE-style policy gradient with a dense, per-token reward. Toggle to see each derivation.

Reverse KL wins Avg@k — but the gap shrinks as horizon grows

Accuracy advantage of reverse KL over forward KL, in Avg@k points, by training horizon.

Entropy Collapse Under Long-Horizon Reverse KL

Mean per-token predictive entropy across training horizon. Forward KL barely moves; reverse KL heads toward zero.

The Trap: Better Warm Start, Worse RL Ceiling

Qwen3-4B teacher, MATH500 accuracy — before and after a GRPO run bolted onto each warm-start checkpoint. The reverse-KL checkpoint starts higher and ends lower.

Fix 1 — Forward-Heavy KL Mixing

A convex combination of forward and reverse KL at each prefix. As the mix shifts toward reverse KL, accuracy gain climbs but entropy stability collapses. The paper lands forward-heavy, not fifty-fifty.

Fix 2 — Entropy-Gated Length Curriculum

Start short. Extend the horizon only while held-out entropy stays above threshold. Stop extending the moment it drops.

Payoff vs Fixed 4096-Token Training

References

1Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation — Zhao, Xin, Fan, Tong, Li, Shen, 2026
2A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning — Ross, Gordon, Bagnell, 2011
3Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks — Bengio, Vinyals, Jaitly, Shazeer, 2015
4Sequence Level Training with Recurrent Neural Networks — Ranzato, Chopra, Auli, Zaremba, 2016
5On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes — Agarwal et al. (Google DeepMind), 2024
6MiniLLM: Knowledge Distillation of Large Language Models — Gu, Dong, Wei, Huang, 2024
7Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models — Luo, Chuang, Wang, Xu, Han, Zhang, Braverman, 2026
8A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) — Ross, Gordon, Bagnell, 2011
9Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting — Chen, Razin, Narasimhan, Chen, 2025

Interactive companion visualization: full paper viz