A 7B teacher — stronger on every benchmark — completely failed to improve a 1.5B student. A weaker teacher, same student, worked. The paper traces this to a token-level "overlap ratio" between teacher and student candidate distributions: benchmark strength and teachability are different axes.
Classic distillation trains on teacher-written text, which never shows the student what to do once it drifts onto its own (possibly wrong) tokens — exposure bias. OPD has the student generate its own rollout and scores every token it actually visited against the teacher's distribution via reverse KL, giving a dense reward with no verifier.
RLVR (reinforcement learning from verifiable rewards) gives one sparse scalar at the end of a response, creating a brutal credit-assignment problem across hundreds of tokens. OPD's teacher log-probabilities give a reward at every token, for free.
In successful runs, overlap ratio climbs from ~72% past 91% while the entropy gap collapses toward zero. In failing runs, both curves sit flat from step one.
Overlap tokens — where teacher and student candidate sets agree — carry 97–99% of total probability mass throughout training. It's the dominant mass shifting into agreement, not long-tail alignment.
Training only on the overlap-token subset nearly matches full top-k performance. Training only on non-overlap tokens is much weaker and even temporarily destabilizes overlap growth.
Matched pairs where initial overlap ratio is nearly identical (~70–75%) across both teachers in a family. The only variable that predicts gap recovery is whether the teacher has RL-acquired capability the student hasn't seen.
JustRL-1.5B (already RL-tuned) distilled backward toward its own pre-RL checkpoint regresses to pre-RL performance. Swapping in R1-Distill-7B — bigger, higher benchmark score — as teacher produces nearly the identical collapse.
Sweeping max response length from 500 to 15K tokens on the same pairing shows a sweet spot around 3K–7K. Too short starves training signal; too long plateaus, then destabilizes — overlap ratio collapses, entropy and gradient norm spike, starting at the end of the response and propagating backward.
Teacher accuracy advantage over a student rollout continued from a truncated prefix: strong near the start, nearly gone by 16K tokens.
Sequence-level mean reward still separates correct from incorrect answers about equally well for both teachers — the "failing" 7B teacher's reward is if anything slightly more globally informative (AUROC), even though its per-token gradient contribution collapses at depth. Unverified hypothesis: anisotropic, direction-canceling per-token advantages.
Off-policy cold start narrows the gap before OPD even begins. Teacher-aligned prompts (template and content) push overlap growth higher, though matching content too tightly can collapse exploration unless mixed with out-of-distribution prompts.