1 00:00:01,000 --> 00:01:03,125 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called "Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation," by Anhao Zhao et al. — six authors total: Anhao Zhao, Haoran Xin, Yingqi Fan, Junlong Tong, Wenjie Li, and Xiaoyu Shen — out of Eastern Institute of Technology in Ningbo, The Hong Kong Polytechnic University, Shanghai Jiao Tong University, and HKUST Guangzhou. Submitted to arXiv on May 16th, 2026. And Ada, here's the thing that got me: every big distillation recipe you can name — DeepSeek-R1, Qwen3, MiMo, GLM-5 — pairs the same two choices together every single time. Off-policy training always comes with forward KL. On-policy training always comes with reverse KL. Nobody seems to have asked whether that pairing is actually necessary, or just habit. 2 00:01:03,125 --> 00:01:44,424 [Dr. Ada Shannon] Right, and that's the part that should bug anyone who's actually implemented one of these pipelines. Because on paper, prefix source — whose text you're training on — and KL direction — which way you're measuring the divergence — are two completely separate knobs. There's no law of math that says student-generated rollouts must be paired with reverse KL, or that teacher traces must be paired with forward KL. But if you go look at actual production recipes, that's exactly what happens, every time, no exceptions. So either there's a deep reason those two choices are locked together, or everyone's just been copying the same two corners of a four-corner design space without checking what's in the other two corners. That's the question this paper actually goes and answers empirically, not just conceptually. 3 00:01:44,424 --> 00:02:01,650 [Hal Turing] Okay so before we get into what's in those other corners, I want to make sure listeners have the vocabulary, because this paper leans hard on a few terms. Let's start basic — knowledge distillation. What is it, actually, at the level of 'you're training a smaller model'? 4 00:02:01,650 --> 00:02:40,400 [Dr. Ada Shannon] Distillation is training a smaller student model to match a bigger teacher model's output distribution, instead of just training the student on hard ground-truth labels the way you'd do standard supervised learning. The idea traces back to Hinton, Vinyals, and Dean's 2015 distillation paper out of Google — the insight there was that a teacher's full probability distribution over next tokens carries more information than a single correct answer does. A wrong-but-plausible guess still tells the student something about the shape of the problem. In the LLM world, that's evolved into a whole design space of exactly how you expose the student to that teacher signal, which is precisely what this paper is mapping out. 5 00:02:40,400 --> 00:02:53,300 [Hal Turing] And that's where on-policy distillation comes in, right? Because I've seen that term thrown around and I want the plain version — what makes something 'on-policy' here versus just normal distillation? 6 00:02:53,300 --> 00:03:38,900 [Dr. Ada Shannon] Yeah, sorry, jumping in because this is the crux of it. Standard off-policy distillation trains the student via cross-entropy on text the teacher generated — that's basically SFT with soft targets. On-policy distillation flips that: the student generates its own rollouts during training, and the teacher scores those student-generated tokens with dense, per-token feedback. So supervision happens on states the student itself actually visits, not states the teacher happened to wander into. That distinction connects straight to exposure bias — if you only ever teacher-force the student on teacher prefixes, it never practices recovering from its own mistakes, and small errors can snowball once it's generating on its own at inference time. DAgger, from Ross, Gordon, and Bagnell back in 2011, is the imitation-learning ancestor of this whole idea — train on the states your own policy actually visits, not just the expert's. 7 00:03:38,900 --> 00:03:57,275 [Hal Turing] So exposure bias is basically: you trained on a nice clean path, but the real world doesn't hand you the same path back. Got it. Now, forward versus reverse KL — that's the other axis, and I know reverse KL gets called 'mode-seeking.' What does that actually mean in practice? 8 00:03:57,275 --> 00:04:34,725 [Dr. Ada Shannon] Reverse KL is KL of student given teacher, averaged over student-sampled outputs, and mode-seeking means it concentrates the student's probability mass on a narrower slice of the teacher's high-probability continuations rather than covering the teacher's full spread. Forward KL does the opposite — it tries to cover the teacher's whole distribution, which is why plain SFT, cross-entropy on fixed teacher text, is literally the Monte Carlo approximation of forward-KL distillation on teacher-generated prefixes. That's not a metaphor, the paper shows it's the same objective in a different clothing. 9 00:04:34,725 --> 00:04:53,275 [Hal Turing] Okay, but here's where I'll push back a little — if forward KL is just SFT and reverse KL is just OPD, isn't 'decoupling' them kind of just... relabeling two things everyone already does? I'm not totally convinced this is a new discovery versus a new vocabulary. 10 00:04:53,275 --> 00:05:15,500 [Dr. Ada Shannon] I actually disagree with you there, Hal. Relabeling would be if they stopped at the two corners everyone already uses. They don't — they cross prefix source with KL direction and get four objectives, and two of those, teacher-prefix reverse KL and student-prefix forward KL, had never been systematically studied before this paper. That's not vocabulary, that's two unexplored quadrants of a design space people assumed didn't exist independently. 11 00:05:15,500 --> 00:05:29,225 [Hal Turing] Fair — I'll grant that the two unstudied quadrants are the actual contribution. I was conflating 'renaming the familiar corners' with 'the whole paper,' and that's not fair to what they're doing with the other two. 12 00:05:29,225 --> 00:06:00,700 [Dr. Ada Shannon] Good, we agree then. And to ground all this before we get into results — their setup is a Qwen3-0.6B-Base student distilling from Qwen3-4B and Qwen3-8B teachers, same model family, so no tokenizer mismatch to worry about. Everything's evaluated on math reasoning — AIME24, AMC23, MATH500, GSM8K — trained on the DeepScaleR dataset. Tightly controlled, same architecture family throughout, which is exactly what you want if you're trying to isolate the effect of just these two axes. 13 00:06:00,700 --> 00:06:12,950 [Hal Turing] Before we get to the fixes — walk me through how all four cells of that grid map onto training paradigms people already recognize. Is there a clean story there, or is it more ad hoc? 14 00:06:12,950 --> 00:06:43,850 [Dr. Ada Shannon] Right — same four-cell grid. Each cell maps onto a classical training regime: prefix source sets the policy regime, off-policy for teacher prefixes, on-policy for student, while KL direction sets the learning objective. Teacher-prefix forward KL is off-policy SFT. Student-prefix forward KL is DAgger-style on-policy SFT. Teacher-prefix reverse KL is an offline-RL-style distillation objective. And student-prefix reverse KL is exactly OPD, reframed as dense-reward on-policy RL. 15 00:06:43,850 --> 00:06:54,825 [Hal Turing] And that reframing — forward KL as SFT, reverse KL as RL — that's not just a vibe, there's an actual gradient-level argument underneath it? 16 00:06:54,825 --> 00:07:29,525 [Dr. Ada Shannon] Yeah, that's Proposition 1. Fix a prefix, take the gradient of forward KL with respect to the student's parameters, and it collapses to exactly the SFT gradient toward the teacher's soft targets. Reverse KL is the more interesting one — its gradient works out to a REINFORCE-style policy gradient, where the reward is the log-ratio of teacher probability to student probability at each token, dense and stop-gradient, computed at every position instead of just at rollout's end. So reverse-KL distillation literally is on-policy RL with a built-in per-token reward. 17 00:07:29,525 --> 00:07:36,050 [Hal Turing] So how'd they actually stress-test all four in practice — that's a lot of runs to keep controlled? 18 00:07:36,050 --> 00:07:57,275 [Dr. Ada Shannon] Tightly — a short 128-token horizon and a long 4096-token horizon, standalone for each objective, then every distilled checkpoint gets a GRPO run bolted on as an RL follow-up. Throughout they track accuracy, mean per-token predictive entropy, and average response length — entropy and length are the canaries for whether the student's distribution is staying healthy. 19 00:07:57,275 --> 00:08:07,550 [Hal Turing] Give me the headline number on the KL-direction side then, since that's where the paper spends most of its ink — what does the accuracy-entropy tradeoff look like? 20 00:08:07,550 --> 00:08:48,500 [Dr. Ada Shannon] Reverse KL wins on Avg@k consistently — plus 2.45 points on average, bigger at the short 128-token horizon, plus 3.68, than at 4096 tokens, plus 1.21. But at 4096 tokens it also drives entropy toward collapse and pushes length toward the generation ceiling. For example, with the Qwen3-4B teacher, a student-prefix reverse-KL warm start enters GRPO around 45% on MATH500 and drops to about 36%. The forward-KL warm start begins lower, around 40%, and climbs to about 45%. The 'better' checkpoint becomes the worse RL start. 21 00:08:48,500 --> 00:09:01,250 [Hal Turing] I want to push on that — isn't lower entropy just the model being more confident because it's more correct? Sharper distribution around the right answer sounds like exactly what you'd want. 22 00:09:01,250 --> 00:09:23,750 [Dr. Ada Shannon] No — Pass@k would move if it were just confident and correct. Instead reverse KL's Pass@k is flat or worse than forward KL's at 4096 tokens — the model got narrower without getting more reliably correct across samples. And the RL numbers settle it: a well-calibrated model doesn't degrade under further training. Forty-five to thirty-six percent isn't confidence, that's lost exploration capacity. 23 00:09:23,750 --> 00:09:33,175 [Hal Turing] Fair — accuracy on one sample versus reliability across samples. Where does prefix source fit in, separate from KL direction? 24 00:09:33,175 --> 00:10:08,950 [Dr. Ada Shannon] Prefix source is more a cost-and-quality story than a stability one — entropy and length are mostly governed by KL direction. Student prefixes beat teacher prefixes under matched steps, plus 1.80 Avg@k, plus 2.11 Pass@k. But teacher prefixes are cheaper, since you can cache the rollouts and teacher logits instead of generating student rollouts online. Under matched FLOPs, cached teacher-prefix training reaches roughly 38 to 40 percent on MATH500 within 15 to 20 thousand cumulative TFLOPs — student prefixes need far more compute for it. 25 00:10:08,950 --> 00:10:13,250 [Hal Turing] And training length compounds whichever direction you picked? 26 00:10:13,250 --> 00:10:36,200 [Dr. Ada Shannon] Going from 128 to 4096 tokens raises Avg@k by 2.56 points and Pass@k by 1.88, but forward KL captures most of that gain, plus 3.80, versus reverse KL's plus 1.32. The price for reverse KL is steeper: at 4096 tokens entropy gets closest to zero and length closest to the 8192-token ceiling. Forward KL barely wobbles. 27 00:10:36,200 --> 00:10:41,575 [Hal Turing] So what do they actually do about that, besides defaulting to forward KL? 28 00:10:41,575 --> 00:10:59,075 [Dr. Ada Shannon] Two fixes, one per tradeoff. For KL direction, they mix forward and reverse KL as a convex combination at each prefix — a forward-heavy mixture, not fifty-fifty, stabilizes entropy and length while keeping almost all of reverse KL's accuracy gain. 29 00:10:59,075 --> 00:11:01,000 [Hal Turing] And the length side? 30 00:11:01,000 --> 00:11:20,775 [Dr. Ada Shannon] That's the entropy-gated curriculum — start short, extend the horizon only while held-out entropy stays above a threshold; if it drops, stop extending. Against fixed 4096-token training, it lifts Avg@k by 3.6 points, Pass@k by up to 5.8, and cuts response length by roughly three times. 31 00:11:20,775 --> 00:12:00,475 [Hal Turing] Okay, roughly 3x shorter responses for a 3.6-point Avg@k gain — that's a genuinely good result. But here's what's nagging me, Ada: the student in every single one of these experiments is Qwen3-0.6B-Base, and the teachers are Qwen3-4B and 8B. Same family, same tokenizer, all pretty modest scale. Does any of this — the accuracy-entropy tradeoff, the forward-heavy KL mixing prescription — actually tell us what happens if you're distilling into something like a 7B or 70B student, or pulling from a much bigger, much more capable teacher? 32 00:12:00,475 --> 00:12:30,400 [Dr. Ada Shannon] Honestly, I don't think we know, and the paper doesn't pretend to. Think about that per-token log-ratio reward we mentioned earlier — at 0.6B-to-4B, that gap is narrow enough that the reward signal is informative without being either trivially zero or wildly noisy. Widen that gap to a 70B teacher and a 7B student, and you could get a much noisier, spikier reward landscape — entropy collapse might happen faster and harder, or the dynamics could shift entirely. Sparse-versus-dense reward behavior is scale-and-gap-dependent territory that nobody's mapped here. 33 00:12:30,400 --> 00:13:03,700 [Hal Turing] And it's not just scale — everything is math with verifiable answers. AIME24, AMC23, MATH500, GSM8K, trained on DeepScaleR. But the paper name-drops DeepSeek-R1, Qwen3, GLM-5 as its motivating pipelines, and those do code, agentic tool use, open-ended generation. Plus there's this shared-tokenizer requirement buried in the setup — student and teacher have to share a vocabulary to avoid artifacts in the token-level KL. That's a real constraint, isn't it, not just a footnote? 34 00:13:03,700 --> 00:13:43,050 [Dr. Ada Shannon] It's a real constraint, and it's the one I'd flag hardest. Most production distillation — pulling a proprietary frontier model's behavior into an open-weight architecture with a totally different tokenizer — can't use this decomposition at all without some kind of alignment or approximation layer, which the paper never touches. And on the domain side: entropy and response length work as diversity proxies here because math has one correct final answer and RL reward is unambiguous. In open-ended generation, a longer or lower-entropy response isn't automatically worse — it might just be a different valid answer. The entropy-gated curriculum's stopping rule could misfire entirely outside verifiable-reward settings. 35 00:13:43,050 --> 00:14:08,800 [Hal Turing] Oh wait, hold on — that actually connects to something that's been bugging me this whole episode. Isn't this whole entropy-collapse-under-reverse-KL story already established? MiniLLM, Yuxian Gu and Li Dong and Furu Wei and Minlie Huang, out of Tsinghua and Microsoft Research, 2024 — that paper already showed reverse KL is mode-seeking and narrows the student's distribution. So how much of this is actually new? 36 00:14:08,800 --> 00:14:41,775 [Dr. Ada Shannon] Some of it is genuinely new, some of it isn't, and I think we should be honest about the split. MiniLLM gave you the mechanism — mode-seeking geometry narrows probability mass. What this paper adds is the gradient-level unification, Proposition 1, showing forward KL is SFT and reverse KL is policy gradient, plus the two off-diagonal cells nobody had built before, plus a concrete engineering fix — the KL mixing ratio and the entropy-gated curriculum — validated with actual numbers. That's not nothing. 37 00:14:41,775 --> 00:15:10,050 [Hal Turing] I actually disagree that the fix is as novel as you're framing it, though. The 2026 OPD length-inflation paper — Feng Luo and collaborators out of Rice and Case Western, reference [28] in the paper — already documents this exact length-blowup-under-long-horizon-reverse-KL failure mode and proposes stabilization strategies. So isn't the entropy-gated curriculum just a re-skin of an already-known fix, wrapped in this paper's four-objective taxonomy? 38 00:15:10,050 --> 00:15:53,950 [Dr. Ada Shannon] No, I'll push back there — the taxonomy is what makes it more than a re-skin. [28] treats length inflation as an isolated OPD pathology to patch. This paper places it inside a decomposition where you can ask 'which of four objectives does this fix apply to,' and they only validate it on one or two cells, which is itself a limitation worth naming. So: the empirical phenomenon, largely confirmatory. The unifying lens and the decision framework for practitioners, that part's real. If you're choosing between off-policy SFT and OPD today, this paper's actual advice is concrete — don't default to pure reverse KL at long horizons, mix in forward KL, gate your length curriculum on held-out entropy. Where it breaks: cross-tokenizer teacher-student pairs, non-math domains, and anything past an 8B student, none of which anyone's tested yet. 39 00:15:53,950 --> 00:16:22,300 [Hal Turing] That's a fair place to land. So to wrap up — this paper's real contribution is turning two folk-recipe defaults into a labeled two-axis design space with a gradient-level explanation underneath, plus two practical knobs, KL mixing and the entropy-gated curriculum, that measurably help at the scale they tested. Whether it survives bigger models, messier domains, and mismatched tokenizers is the open question. Ada, thanks as always for keeping me honest on this one. 40 00:16:22,300 --> 00:16:26,000 [Dr. Ada Shannon] Anytime, Hal. That's all for this one — thanks for listening.