Token Teachability: Rethinking Disagreement in On-Policy Distillation
Raw KL divergence tells you two models disagree — it can't tell you whether the student can actually act on the correction. This paper splits "disagreement" into learnable and incompatible components and trains on just 5% of tokens picked by teachability, not magnitude.
arXiv:2605.26844Wang, Lu, Gu, Wang, Yang, Yan, Xie, Wu, Yang · 2026HK PolyU · InfiX.ai · Daya Bay Institute
Lineage: from dark knowledge to on-policy correction
a decade of distillation, one branch at a time
The On-Policy Distillation + TA-OPD pipeline
where teachability filtering slots into the loop
The teacher must still score every token to compute D‑t and C‑t — the 5% figure only shrinks the student's backward pass. More on that in Limits.
Same raw KL, two different stories
toggle between the two disagreement types from Figure 1
Student p(token)Teacher p(token)
Raw KL (D‑t, magnitude)
0.83
Compatibility (C‑t)
0.91
Learnable disagreement (D‑L = D‑t · C‑t)
0.76
Incompatible disagreement (D‑I = D‑t · (1−C‑t))
0.07
In the learnable case, the teacher's correction reweights candidates the student was already entertaining — a small, actionable nudge. Raw KL is high, but so is compatibility.
TIP's confident-disagreement quadrant is not one thing
entropy vs. raw KL, colored by teachability — click legend to isolate a group
Q3 = low entropy, high KL — the "confident but wrong" zone TIP treats as uniformly golden. Splitting it by teachability shows only the green cluster earns its keep; gray is flat, red is actively wasted.
Downstream benchmarks & the budget sweep
average score by teacher→student pair
The messier the pairing, the bigger TA-OPD's edge: on the cross-backbone DeepSeek-R1-Distill-14B→Qwen2.5-3B pair, Full OPD actually dips below the untouched base model while TA-OPD stays positive. No paired significance test is reported on these averages — see Limits.
score vs. retained-token budget (%)
Not monotonic, and the two curves don't even agree on direction: for the 14B teacher, 30% budget dips below the 10% score before recovering slightly at 50%. Best selector and best budget are both pair-specific.
The 5% figure is a supervision budget, not a speedup
teacher cost breakdown: Full OPD vs. TA-OPD
Computing D‑t and C‑t requires the teacher's top-K log-probs at every position, before the mask is applied. The teacher's forward pass never shrinks — only the student's backward pass does.
A dial, not a switch
regression coefficients on downstream gain
Both coefficients are positive — learnable disagreement's is only about 2× incompatible disagreement's. "Incompatible" tokens still carry some signal.
Cross-task inconsistency
TA-OPD − Full OPD, by benchmark and pair
HumanEval and IFEval swing in both directions across pairs, on a model trained entirely on math prompts — the paper doesn't isolate whether that's teachability or generic perturbation.
References
Wang, Lu, Gu, Wang, Yang, Yan, Xie, Wu, Yang. Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation. 2026. arXiv:2605.26844
Hinton, Vinyals, Dean. Distilling the Knowledge in a Neural Network. 2015. scholar link
Kim, Rush. Sequence-Level Knowledge Distillation. 2016. scholar link
Agarwal, Vieillard, et al. (Google DeepMind). On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes. 2024. scholar link
Gu, Dong, Wei, Huang. MiniLLM: Knowledge Distillation of Large Language Models. 2024. scholar link
Ross, Gordon, Bagnell. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger). 2011. scholar link
Lin, Gou, Gong, et al. Not All Tokens Are What You Need for Pretraining (Rho-1). 2024. scholar link
Li, Holtzman, Fried, et al. Contrastive Decoding: Open-ended Text Generation as Optimization. 2023. scholar link
Ko, Kim, Chen, Yun. DistiLLM: Towards Streamlined Distillation for Large Language Models. 2024. scholar link
Leviathan, Kalman, Matias. Fast Inference from Transformers via Speculative Decoding. 2023. scholar link
Xu, Sang, Zhou, et al. TIP: Token Importance in On-Policy Distillation. 2026. scholar link
Jin, Min, Yang, et al. Entropy-Aware On-Policy Distillation of Language Models. 2026. scholar link
Li, Zuo, He, et al. Rethinking On-Policy Distillation of Large Language Models. 2026. scholar link
Wang, Yu, Gao, et al. Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective RL for LLM Reasoning. 2026. scholar link
Guo, Yang, Zhang, et al. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL. 2025. scholar link