arXiv:2604.13016 Tsinghua · ShanghaiTech · UIUC · Renmin Posted April 15, 2026

On-Policy Distillation: Why a Stronger Teacher Can Backfire

AI Post Transformers · Hal Turing & Dr. Ada Shannon

A 7B teacher — stronger on every benchmark — completely failed to improve a 1.5B student. A weaker teacher, same student, worked. The paper traces this to a token-level "overlap ratio" between teacher and student candidate distributions: benchmark strength and teachability are different axes.

Where does the reward signal come from? Classic SFT distillation vs. on-policy distillation

Classic distillation trains on teacher-written text, which never shows the student what to do once it drifts onto its own (possibly wrong) tokens — exposure bias. OPD has the student generate its own rollout and scores every token it actually visited against the teacher's distribution via reverse KL, giving a dense reward with no verifier.

Reward density: RLVR vs OPD

RLVR (reinforcement learning from verifiable rewards) gives one sparse scalar at the end of a response, creating a brutal credit-assignment problem across hundreds of tokens. OPD's teacher log-probabilities give a reward at every token, for free.

Progressive alignment: overlap ratio & entropy gap over training

In successful runs, overlap ratio climbs from ~72% past 91% while the entropy gap collapses toward zero. In failing runs, both curves sit flat from step one.

Successful run Failing run

Token-level overlap (hover a token)

Overlap tokens — where teacher and student candidate sets agree — carry 97–99% of total probability mass throughout training. It's the dominant mass shifting into agreement, not long-tail alignment.

Ablation: train on overlap tokens only

Training only on the overlap-token subset nearly matches full top-k performance. Training only on non-overlap tokens is much weaker and even temporarily destabilizes overlap growth.

Same overlap, different outcome: RL-acquired knowledge is what transfers

Matched pairs where initial overlap ratio is nearly identical (~70–75%) across both teachers in a family. The only variable that predicts gap recovery is whether the teacher has RL-acquired capability the student hasn't seen.

5.3% → 16.9%
DeepSeek: same-pipeline vs RL-tuned teacher
15.6% → 58.6%
Qwen3: non-thinking vs RL-Math teacher

Reverse-distillation: a stronger teacher can't out-run a compatible pattern

JustRL-1.5B (already RL-tuned) distilled backward toward its own pre-RL checkpoint regresses to pre-RL performance. Swapping in R1-Distill-7B — bigger, higher benchmark score — as teacher produces nearly the identical collapse.

JustRL-1.5B (RL-tuned baseline) Distilled toward pre-RL self Distilled toward R1-Distill-7B

Response length sweep: the reward degrades over long horizons

Sweeping max response length from 500 to 15K tokens on the same pairing shows a sweet spot around 3K–7K. Too short starves training signal; too long plateaus, then destabilizes — overlap ratio collapses, entropy and gradient norm spike, starting at the end of the response and propagating backward.

Teacher advantage collapses with prefix depth

Teacher accuracy advantage over a student rollout continued from a truncated prefix: strong near the start, nearly gone by 16K tokens.

Informative globally, useless locally

Sequence-level mean reward still separates correct from incorrect answers about equally well for both teachers — the "failing" 7B teacher's reward is if anything slightly more globally informative (AUROC), even though its per-token gradient contribution collapses at depth. Unverified hypothesis: anisotropic, direction-canceling per-token advantages.

The recipe: two levers for a mismatched pairing

Off-policy cold start narrows the gap before OPD even begins. Teacher-aligned prompts (template and content) push overlap growth higher, though matching content too tightly can collapse exploration unless mixed with out-of-distribution prompts.

References