AI Post Transformers · Episode Companion

Token Teachability: Rethinking Disagreement in On-Policy Distillation

Raw KL divergence tells you two models disagree — it can't tell you whether the student can actually act on the correction. This paper splits "disagreement" into learnable and incompatible components and trains on just 5% of tokens picked by teachability, not magnitude.

arXiv:2605.26844 Wang, Lu, Gu, Wang, Yang, Yan, Xie, Wu, Yang · 2026 HK PolyU · InfiX.ai · Daya Bay Institute

Lineage: from dark knowledge to on-policy correction

a decade of distillation, one branch at a time

The On-Policy Distillation + TA-OPD pipeline

where teachability filtering slots into the loop

The teacher must still score every token to compute D‑t and C‑t — the 5% figure only shrinks the student's backward pass. More on that in Limits.

Same raw KL, two different stories

toggle between the two disagreement types from Figure 1

Student p(token) Teacher p(token)
Raw KL (D‑t, magnitude)
0.83
Compatibility (C‑t)
0.91
Learnable disagreement (D‑L = D‑t · C‑t)
0.76
Incompatible disagreement (D‑I = D‑t · (1−C‑t))
0.07

In the learnable case, the teacher's correction reweights candidates the student was already entertaining — a small, actionable nudge. Raw KL is high, but so is compatibility.

TIP's confident-disagreement quadrant is not one thing

entropy vs. raw KL, colored by teachability — click legend to isolate a group

Q3 = low entropy, high KL — the "confident but wrong" zone TIP treats as uniformly golden. Splitting it by teachability shows only the green cluster earns its keep; gray is flat, red is actively wasted.

Downstream benchmarks & the budget sweep

average score by teacher→student pair

The messier the pairing, the bigger TA-OPD's edge: on the cross-backbone DeepSeek-R1-Distill-14B→Qwen2.5-3B pair, Full OPD actually dips below the untouched base model while TA-OPD stays positive. No paired significance test is reported on these averages — see Limits.

The 5% figure is a supervision budget, not a speedup

teacher cost breakdown: Full OPD vs. TA-OPD

Computing D‑t and C‑t requires the teacher's top-K log-probs at every position, before the mask is applied. The teacher's forward pass never shrinks — only the student's backward pass does.

A dial, not a switch

regression coefficients on downstream gain

Both coefficients are positive — learnable disagreement's is only about 2× incompatible disagreement's. "Incompatible" tokens still carry some signal.

Cross-task inconsistency

TA-OPD − Full OPD, by benchmark and pair

HumanEval and IFEval swing in both directions across pairs, on a model trained entirely on math prompts — the paper doesn't isolate whether that's teachability or generic perturbation.

References

  1. Wang, Lu, Gu, Wang, Yang, Yan, Xie, Wu, Yang. Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation. 2026. arXiv:2605.26844
  2. Hinton, Vinyals, Dean. Distilling the Knowledge in a Neural Network. 2015. scholar link
  3. Kim, Rush. Sequence-Level Knowledge Distillation. 2016. scholar link
  4. Agarwal, Vieillard, et al. (Google DeepMind). On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes. 2024. scholar link
  5. Gu, Dong, Wei, Huang. MiniLLM: Knowledge Distillation of Large Language Models. 2024. scholar link
  6. Ross, Gordon, Bagnell. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger). 2011. scholar link
  7. Lin, Gou, Gong, et al. Not All Tokens Are What You Need for Pretraining (Rho-1). 2024. scholar link
  8. Li, Holtzman, Fried, et al. Contrastive Decoding: Open-ended Text Generation as Optimization. 2023. scholar link
  9. Ko, Kim, Chen, Yun. DistiLLM: Towards Streamlined Distillation for Large Language Models. 2024. scholar link
  10. Leviathan, Kalman, Matias. Fast Inference from Transformers via Speculative Decoding. 2023. scholar link
  11. Xu, Sang, Zhou, et al. TIP: Token Importance in On-Policy Distillation. 2026. scholar link
  12. Jin, Min, Yang, et al. Entropy-Aware On-Policy Distillation of Language Models. 2026. scholar link
  13. Li, Zuo, He, et al. Rethinking On-Policy Distillation of Large Language Models. 2026. scholar link
  14. Wang, Yu, Gao, et al. Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective RL for LLM Reasoning. 2026. scholar link
  15. Guo, Yang, Zhang, et al. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL. 2025. scholar link