1 00:00:01,000 --> 00:00:58,674 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called "Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe." That's Yaxuan Li et al., eleven authors total, out of Tsinghua University, ShanghaiTech University, University of Illinois Urbana-Champaign, and Renmin University of China, posted to arXiv on April 15th, 2026. And Ada, here's the number that stopped me cold: they took a 7-billion-parameter teacher — stronger by every benchmark you'd check — used it to distill a 1.5-billion-parameter student, and the run completely failed. Then they swapped in a smaller, weaker teacher, same student, and it worked great. That's the kind of result that should make anyone running a distillation pipeline right now nervous. 2 00:00:58,674 --> 00:01:38,349 [Dr. Ada Shannon] It should, because it breaks an assumption almost everybody in post-training has been quietly operating on: that if you upgrade the teacher, the student gets better, full stop. This paper is basically a systematic dismantling of that assumption for a specific, increasingly popular technique called on-policy distillation, or OPD. The core question they're chasing is exactly what you just described — under what conditions does OPD actually work, and what's happening at the token level when it does or doesn't? Before we get into their answer, though, we should back up, because OPD itself is relatively new to this show, and it's worth being precise about what makes it different from plain old distillation. 3 00:01:38,349 --> 00:02:01,500 [Hal Turing] Right, let's do that. So classic knowledge distillation — the stuff people have been doing for years — you take a big teacher model, have it generate a bunch of text, and then fine-tune your smaller student on those teacher-written sequences with normal next-token cross-entropy loss. It's just supervised fine-tuning on synthetic data. What's wrong with that? It sounds reasonable. 4 00:02:01,500 --> 00:02:44,400 [Dr. Ada Shannon] It sounds reasonable until you remember imitation learning has been fighting this exact problem since at least Ross, Gordon, and Bagnell's 2011 DAgger paper. The issue is called exposure bias, and Bengio and colleagues named it explicitly back in 2015 for sequence models: the student only ever sees teacher-written prefixes during training. It never learns what to do once it drifts off that path and starts looking at its own, possibly wrong, tokens as context. So at inference time, one early mistake compounds, because the student was never trained to recover from a state it put itself in. On-policy distillation flips who generates the trajectories — the student samples its own rollouts, exactly like it would at inference, and then at every token position it actually visited, you query the teacher for its probability distribution over what comes next. 5 00:02:44,400 --> 00:02:54,200 [Hal Turing] So the reward is just — how far off is the student's guess from the teacher's guess, at every single step, on the student's own generated path. 6 00:02:54,200 --> 00:03:47,375 [Dr. Ada Shannon] Exactly, and that difference is typically measured as reverse KL divergence between the student and teacher distributions. Reverse KL matters here because it's mode-seeking rather than mode-covering — it pushes the student to concentrate probability mass on the teacher's high-confidence modes rather than spreading itself thin trying to cover everything the teacher considers merely plausible. Compare that to outcome-reward RL, what people call RLVR, reinforcement learning from verifiable rewards — checking whether a final math answer or a bit of code is correct. RLVR also has the student generate its own rollouts, but the reward is one sparse scalar at the very end of potentially hundreds of tokens, which creates a brutal credit-assignment problem. OPD sidesteps that entirely because the teacher's log-probabilities give you a dense reward at every token, for free, no verifier or reward model required. 7 00:03:47,375 --> 00:04:00,849 [Hal Turing] Oh wait wait wait — so is that basically the pitch, then? Bigger, stronger teacher just means a stronger reward signal at every one of those steps, so obviously scaling up the teacher should only help? 8 00:04:00,849 --> 00:04:31,949 [Dr. Ada Shannon] I actually disagree with the way you're framing that, Hal, and that's precisely the trap this whole paper is built to expose. A stronger teacher by benchmark score doesn't automatically mean a more useful reward signal for this particular student. If the teacher's top-k next-token candidates at a given state barely overlap with the student's own candidates — what the authors call the overlap ratio — then the teacher is effectively speaking a different dialect, and the dense signal becomes dense noise. 9 00:04:31,949 --> 00:04:39,425 [Hal Turing] Hang on, but a higher benchmark score has to mean something, doesn't it? I'm not ready to just throw the score out. 10 00:04:39,425 --> 00:05:18,000 [Dr. Ada Shannon] It means something about the teacher in isolation. It doesn't mean the teacher has anything new to teach this student. That's the whole distinction the paper is built around, and honestly, I think that's the more interesting half of their argument — score and teachability are just not the same axis. We'll get into exactly how they show that. For now, the structure of the paper is three parts: phenomenology, where they document when OPD succeeds or fails; mechanism, where they trace it down to this token-level overlap behavior; and recipe, where they propose fixes for the failure cases. The headline result you opened with is their clearest phenomenology finding. 11 00:05:18,000 --> 00:05:27,649 [Hal Turing] So let's get concrete, Ada. What did they actually run to test whether thinking-pattern compatibility matters more than raw teacher strength? 12 00:05:27,649 --> 00:06:09,549 [Dr. Ada Shannon] They took Qwen3-1.7B-Base as the student and pitted two teachers against each other: Qwen3-4B in its non-thinking mode, and a version of Qwen3-4B-Base with GRPO applied as a zero-RL run. On raw benchmark scores the two teachers are basically neck and neck, the GRPO teacher even loses on AMC 2023, 59.9 versus 70 percent. But the GRPO teacher's overlap ratio with the student starts noticeably higher, and distillation from it consistently outperforms the non-thinking teacher across training. The overlap curves eventually converge, but the accuracy gap never closes. Whatever gets lost in that early mismatch window doesn't come back. 13 00:06:09,549 --> 00:06:16,974 [Hal Turing] Okay, so that's condition one. What's the setup for condition two, the higher-score-doesn't-mean-new-knowledge piece? 14 00:06:16,974 --> 00:07:01,349 [Dr. Ada Shannon] Two matched pairs. In the DeepSeek family, R1-Distill-1.5B as student, comparing R1-Distill-7B, same pipeline just bigger, against Skywork-OR1-Math-7B, which is R1-Distill-7B after further RL. Gap recovery for the same-pipeline teacher is 5.3 percent. For the RL-post-trained teacher, 16.9 percent. Same story in the Qwen family: Qwen3-4B non-thinking recovers only 15.6 percent, but Qwen3-4B-Non-Thinking-RL-Math recovers 58.6 percent. Here's the kicker: initial overlap ratios in both pairs sit around 70 to 75 percent, nearly identical. It's not a thinking-pattern story here. The only variable that predicts outcome is whether the teacher has RL-acquired capability the student hasn't seen. 15 00:07:01,349 --> 00:07:13,349 [Hal Turing] Oh wait wait wait, hold on, they didn't just leave it at correlation, did they? Because 'same overlap, different outcome' is suggestive but I'd want something that actually isolates it. 16 00:07:13,349 --> 00:07:48,074 [Dr. Ada Shannon] Right, that's the reverse-distillation experiment, and it's the cleanest result in the paper. Take JustRL-1.5B, which is R1-Distill-1.5B after RL, and distill it backward toward its own pre-RL checkpoint. It regresses almost exactly to pre-RL performance, erasing the RL gains. Swap in R1-Distill-7B as teacher instead, a bigger model scoring slightly higher than JustRL-1.5B. Same regression, nearly indistinguishable trajectory. A higher-scoring teacher produced the identical collapse as the student's own weaker past self, because from the student's rollout distribution, they aren't distinguishable targets. 17 00:07:48,074 --> 00:08:01,449 [Hal Turing] Honestly that makes me think OPD isn't teaching the student anything new most of the time, it's just stamping the teacher's style over whatever the student already does, and we're calling style transfer 'learning.' 18 00:08:01,449 --> 00:08:35,899 [Dr. Ada Shannon] I'd push back hard on that, Hal. The reverse-distillation result shows the student's own pattern gets overwritten, sure, but the Skywork and Qwen3-RL-Math comparisons show real capability transfer too, 58.6 percent gap recovery isn't stylistic mimicry, that's the student inheriting math competence it didn't have. OPD does two different things depending on the teacher: compatible pattern with no new knowledge gets you overwriting with no gain, like the reverse case. Genuine new knowledge behind a compatible pattern gets you real transfer. Collapsing those into 'it's all style' throws away the one result that matters for picking a teacher. 19 00:08:35,899 --> 00:08:45,199 [Hal Turing] Fair, I'll take that. So walk me through what's actually happening token by token when it works, what does 'progressive alignment' look like on a chart? 20 00:08:45,199 --> 00:09:28,549 [Dr. Ada Shannon] In successful runs, overlap ratio climbs steadily from around 72 percent up past 91, the entropy gap narrows toward zero, and, this is the detail I like, those overlap tokens carry 97 to 99 percent of the total probability mass the entire time. It's the dominant mass shifting into agreement, not some long-tail alignment. In failing runs, all three curves sit flat from step one. Then they ran the ablation to check causation: train only on the overlap-token subset, and you match full top-k performance almost exactly. Train only on non-overlap tokens, and it's much weaker and even destabilizes overlap temporarily. It's self-reinforcing, once a token lands in the shared high-probability set, reverse-KL updates keep piling mass onto it and squeeze out the competitors. 21 00:09:28,549 --> 00:09:34,849 [Hal Turing] So if you're stuck with a mismatched teacher, is there anything short of picking a different one? 22 00:09:34,849 --> 00:10:18,549 [Dr. Ada Shannon] Two things, attacking the gap from different ends. First, off-policy cold start: SFT the student on 200K teacher rollouts before OPD ever starts, they did this with Qwen3-1.7B-Base and Qwen3-4B non-thinking, and the SFT-initialized student starts with much higher overlap and a smaller entropy gap that persists the whole way through training. Second, teacher-aligned prompts, at two levels. Swapping just the prompt template to match what the teacher saw in post-training boosts both accuracy and overlap growth. Matching prompt content, using the teacher's actual RL data instead of merely in-domain data, concentrates student mass more tightly on fewer overlap tokens, but tanks student entropy enough that they recommend mixing in out-of-distribution prompts to avoid collapsing exploration entirely. 23 00:10:18,549 --> 00:10:36,199 [Hal Turing] Okay, so those two levers can rescue a mismatched pairing. But Section 6 turns around and complicates the whole picture, because they go digging into whether that dense per-token reward is even reliable, and the answer depends a lot on how long the response is. 24 00:10:36,199 --> 00:11:18,924 [Dr. Ada Shannon] Right, they sweep max response length from 500 tokens up to 15K on the same pairing — R1-Distill-1.5B against JustRL-1.5B — and there's a clear sweet spot around 3K to 7K tokens. Too short and there aren't enough supervised tokens to learn from efficiently. Too long and performance plateaus then drops, with visible instability — overlap ratio collapses, entropy and gradient norm spike. The instability isn't uniform either — it starts at the end of the response and propagates backward toward earlier tokens as training goes on. And when they test whether the teacher can still improve a student rollout from a truncated prefix, its accuracy advantage collapses — plus 0.37 at a 1K-token prefix, down to plus 0.02 at 16K. 25 00:11:18,924 --> 00:11:35,125 [Hal Turing] Oh wait wait wait — plus 0.02? That's basically nothing, the teacher barely does better than the student it's supposed to be correcting. Does the reward signal just become garbage that deep into a trajectory, or is something else going on? 26 00:11:35,125 --> 00:12:18,475 [Dr. Ada Shannon] That's the interesting part — it's not garbage globally. They compute a sequence-level mean reward per rollout and check whether it separates correct from incorrect answers, and both teachers do that reliably — JustRL-1.5B gets AUROC 0.73, R1-Distill-7B gets 0.75, so the failing teacher's reward is if anything slightly more globally informative. Their explanation is that the 7B teacher's per-token advantages might be anisotropic — large individually but pointing in inconsistent directions across the sequence, so aggregating them into a gradient partially cancels out. Informative globally, useless locally. Credit to them, though — they say outright this is unverified, no gradient-direction analysis, it's explicitly a hypothesis for future work. 27 00:12:18,475 --> 00:12:55,725 [Hal Turing] Credit to them for saying so instead of dressing it up as a finding. But that honesty points at a bigger issue I have with this paper's scope. Every experiment here — phenomenology, mechanism, recipe, all of it — is math. AIME, AMC, DAPO-Math prompts. And the model pairs are basically two lineages, Qwen3 and DeepSeek-R1-Distill plus its RL descendants, same tokenizer, same architecture, overlapping pretraining almost by construction. So when they say 'thinking-pattern consistency governs OPD,' how much of that is really just 'these models share a family tree'? 28 00:12:55,725 --> 00:13:54,975 [Dr. Ada Shannon] I'd push back on part of that, Hal. Within what they tested, the evidence is tight — the reverse-distillation result especially, where R1-Distill-7B and R1-Distill-1.5B drive JustRL-1.5B to the exact same regressed floor, that's a controlled, matched comparison, not hand-waving. Where I'll agree is the title — 'Rethinking On-Policy Distillation of Large Language Models' reads like a general theory, and their own future-work section admits cross-family tests confound tokenizer and architecture with genuine data divergence, and that controlled pretraining ablations are, quote, prohibitively expensive. It's not entirely new territory either — Cho and Hariharan showed capacity gaps hurt distillation back in 2019, and Busbridge and colleagues at Apple formalized it as a U-shaped scaling law in 2025. What's new is pushing that down to the token level, building on the reverse-KL, mode-seeking framing MiniLLM set up — Gu, Dong, Wei, and Huang, 2023. 29 00:13:54,975 --> 00:14:10,925 [Hal Turing] Fair — so it's less 'new phenomenon' and more 'known capacity-gap problem, now with a token-level microscope on it.' I can live with that framing. So practically, if I'm building an OPD pipeline tomorrow, what do I actually do differently? 30 00:14:10,925 --> 00:14:44,675 [Dr. Ada Shannon] Pick your teacher for overlap and lineage compatibility, not leaderboard score — a bigger same-family model that just fits the same data harder won't teach you much. Stuck with a mismatch? Cold-start SFT first. And don't push response length past what the reward can track — 15K-token math CoT already shows cracks, so long-horizon or agentic settings with interleaved environment feedback are genuinely open territory the paper doesn't touch. They flag that themselves, along with self-distillation, where thinking-pattern consistency is free by construction and only the new-knowledge condition is doing any work. 31 00:14:44,675 --> 00:15:08,075 [Hal Turing] Good note to end on. Bottom line: a stronger teacher isn't automatically a better one — what matters is whether it shares your student's thinking patterns and actually knows something the student doesn't. That holds up well within math reasoning on closely related model families, even if the title promises more than that. Thanks for listening, everyone — we'll catch you next time.