1 00:00:01,000 --> 00:00:30,899 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called Weak-to-Strong On-Policy Distillation, from Fangxu Yu et al., four co-authors, out of University of Maryland College Park, Microsoft Research, and MBZUAI. And Ada, the number that jumped out at me right away: their student model actually beats the teacher it learned from. Not matches — beats. 2 00:00:30,899 --> 00:01:05,299 [Dr. Ada Shannon] Right, and normally that sentence shouldn't parse. Distillation is supposed to be a ceiling operation — the student chases the teacher's distribution, so how do you end up above the thing you're chasing? That's the tension this whole paper hangs on. Every teacher they use here is weaker than the 8-billion-parameter student doing the learning, and on the math benchmarks the student still comes out ahead of its own domain expert. So the question isn't 'can a small model teach a big one anything,' it's 'why does this particular recipe let weak supervision produce strong-model gains instead of just capping the student at the teacher's level.' 3 00:01:05,299 --> 00:01:22,299 [Hal Turing] So let's set the stage, because there are a couple of ingredients here our listeners have heard us mention before but maybe not unpacked fully. Ada, walk me through on-policy distillation, because I know it's not the classic Hinton-style distillation from 2015. 4 00:01:22,299 --> 00:02:19,800 [Dr. Ada Shannon] Right, so classic knowledge distillation — Hinton, Vinyals, and Dean, 2015 — trains a small student to match a big teacher's softened output probabilities over a fixed dataset. It's offline: the training text comes from wherever the dataset came from, the student never generated it. That's fine for image classifiers, but for an LLM producing long chains of tokens, it creates exposure bias — a term formalized by Ross, Gordon, and Bagnell's DAgger paper in 2011. So the first small mistake mid-generation lands the student in a state nobody trained it to recover from, and errors compound from there. On-policy distillation flips the source: the student generates its own rollouts, the teacher scores every token of those rollouts, and the student minimizes per-token reverse KL against the teacher's distribution on its own trajectories. It's always practicing recovery from exactly the states it actually wanders into. 5 00:02:19,800 --> 00:02:32,325 [Hal Turing] Okay, and that's different from just doing reinforcement learning with a verifiable reward, right? Because I think people hear 'the model learns from its own generations' and assume that's RL. 6 00:02:32,325 --> 00:03:12,500 [Dr. Ada Shannon] Good instinct to check, because they sit right next to each other. RLVR — reinforcement learning from verifiable rewards, the GRPO and PPO-style training behind most reasoning models — is fully on-policy too, but it only gives you one scalar reward at the end of an entire rollout. Hundreds of tokens, one pass-fail signal, so credit assignment is noisy. On-policy distillation keeps the on-policy rollout collection but swaps that sparse scalar for the teacher's dense per-token log-probabilities. You get on-policy data and dense supervision at the same time — which is exactly why the paper frames RLVR and classic distillation as complementary but each incomplete on its own. 7 00:03:12,500 --> 00:03:29,000 [Hal Turing] Wait, hold on — so before we go further, I actually want to push on something. Isn't this just DAgger with extra vocabulary? Like, 'query something better on the states you actually visit' is the same idea Ross and Bagnell had in 2011. 8 00:03:29,000 --> 00:03:50,724 [Dr. Ada Shannon] It's the same lineage, sure, but I wouldn't wave it off as relabeling. DAgger queries an expert that's assumed to be correct. The entire premise of this paper is that at the frontier there is no expert — the 'teacher' is weaker than your student by construction. So the mechanism is DAgger-shaped, but the supervision-quality assumption underneath it is inverted, and that inversion is the actual research contribution, not the rollout-collection idea. 9 00:03:50,724 --> 00:04:05,474 [Hal Turing] Okay, that's fair — I'll grant the mechanism is old but the setting is genuinely different. So let's name that setting properly: this is weak-to-strong generalization, which listeners might remember from OpenAI's alignment work. 10 00:04:05,474 --> 00:04:48,024 [Dr. Ada Shannon] Burns, Izmailov, and a large OpenAI Superalignment author list, 2023. Their question was almost philosophical: finetune a strong pretrained model on labels from a much weaker, often-wrong supervisor — does it just inherit the weak model's error rate? They found it doesn't — it partially generalizes past the weak supervisor's mistakes, recovering some of the gap to ground truth, because its own pretrained knowledge acts as a prior. That was proposed as a stand-in for a future problem: once models are more capable than any available human or supervisor, how do you even produce training signal for them? This W2S-OPD paper borrows that exact framing but points it at capability, not alignment — can you improve a strong model's actual reasoning using signal built from weaker models, rather than teach it to be honest despite bad labels. 11 00:04:48,024 --> 00:05:16,099 [Hal Turing] And that framing matters because of where the industry actually is right now. Distillation has been used at scale — Qwen3's technical report shrinks their 235-billion mixture-of-experts model into smaller dense checkpoints, and multi-teacher on-policy distillation, MOPD, trains several domain-expert teachers from a shared base and consolidates them into one student. But both of those still assume a teacher at least as capable as the student somewhere in the pipeline. 12 00:05:16,099 --> 00:05:49,449 [Dr. Ada Shannon] Which is exactly the wall. Strong-to-weak distillation collapses once there's no larger teacher to shrink. MOPD sidesteps needing a stronger teacher, but it pays for that by requiring costly training at the student's own scale for every domain expert. Both approaches assume you can afford a teacher at or above the student's level. This paper's move is to drop that assumption entirely and ask whether a strong student can keep improving using nothing but a handful of models that are already weaker than it — which is the setup we'll get into next. 13 00:05:49,449 --> 00:06:09,324 [Hal Turing] Right, at least as capable as the student. So walk me through the actual trick, Ada, because this is the part I couldn't picture from the abstract. If every source model is weaker than your 8-billion-parameter student, how do you build something that functions as a teacher at all? What's literally happening to the numbers? 14 00:06:09,324 --> 00:06:39,675 [Dr. Ada Shannon] It's a logit-space construction, and it's cleaner than it sounds. You take a positive model and a negative model — both small — run them on the student's own generated text, and subtract the negative model's logits from the positive model's logits at every token. That difference is Equation 2 in the paper: you add it, scaled by an amplification coefficient alpha, onto the student's own base model's logits, then softmax the result. So the 'teacher' is literally the student's own distribution plus an injected direction, not some external model's distribution. 15 00:06:39,675 --> 00:06:48,150 [Hal Turing] Okay, so why go through all that instead of just distilling straight from the stronger of the two weak models? Wouldn't that be simpler? 16 00:06:48,150 --> 00:07:20,225 [Dr. Ada Shannon] Simpler, and worse. Two things happen when you subtract. First, whatever both weak models get wrong in common — their shared 'ability ceiling' — cancels out in the subtraction, so you're not importing their mediocrity, just the delta between them. Second, because you're adding that delta onto the student's own base logits rather than replacing them, the resulting proxy teacher stays distributionally close to what the student already produces. Direct distillation from a weak model drags the student toward the weak model's whole distribution, errors included. This only injects the direction of improvement. 17 00:07:20,225 --> 00:07:27,750 [Hal Turing] And they don't just do this one way — there are three separate recipes for picking that positive-negative pair, right? 18 00:07:27,750 --> 00:08:00,800 [Dr. Ada Shannon] Three, and each isolates something different. First: a post-RL domain expert as positive, its own pre-RL checkpoint as negative — that isolates whatever skill the RL run instilled, without ever running RL at the student's scale. Second: a larger off-the-shelf base model against a smaller one, same family — that isolates the capability that comes purely from scale, for free, no training at all. Third: one base model conditioned on a correct solution hint versus a wrong one — that isolates an instance-level 'direction toward the right answer' rather than a general skill. 19 00:08:00,800 --> 00:08:14,725 [Hal Turing] Oh wait, wait — hold on, that's actually the part I wanted to ask about. If you've got several of these positive-negative pairs lying around, can you just add all their directions onto the same base model at once? 20 00:08:14,725 --> 00:08:35,850 [Dr. Ada Shannon] Exactly right, and that's Equation 4. You sum multiple weighted capability directions — math expert minus its pre-RL init, code expert minus its pre-RL init, whatever you've got — onto one shared base model, and distill all of it in a single OPD run. One student absorbs several domain skills without training separate consolidation passes the way MOPD requires. 21 00:08:35,850 --> 00:08:39,925 [Hal Turing] So what did they actually run this on, and did it hold up? 22 00:08:39,925 --> 00:09:41,200 [Dr. Ada Shannon] Qwen3-8B as the student, with Qwen3-4B, a Qwen3-4B-RL checkpoint, and Qwen3-0.6B as the weak sources. Four math benchmarks — AIME24, AIME25, both HMMT25 splits — plus HumanEval+, MBPP+, and LiveCodeBench for code. The pre/post-RL setting wins biggest: math average of 51.8 versus OPD's 46.5, and it actually clears the 4B-RL teacher's own 48.8. The scale contrast and the hint contrast still lift the student too, even though every single source model involved is weaker than the 8B student itself. Alpha has a sweet spot — too small and you barely inject the direction, too large and the proxy teacher drifts too far from the student's own distribution and performance drops. And trained only on math, this method actually improves GPQA-Diamond and IFBench, whereas plain OPD can drag general ability below the untouched base model. 23 00:09:41,200 --> 00:09:55,900 [Hal Turing] That out-of-domain result is the one that surprises me most, honestly — you'd expect distilling hard on math to cost you something elsewhere. Did they look at what's different token-by-token between the three contrasts? 24 00:09:55,900 --> 00:10:36,050 [Dr. Ada Shannon] They did, using an episode classifier over Schoenfeld's problem-solving taxonomy. The post-RL and hint contrasts load onto Plan and Monitor tokens — the scaffolding that structures and tracks the solution. The scale contrast instead loads onto Analyze and Implement — the actual solving mechanics. So RL and hinting teach the framework, raw scale teaches execution. That's a genuinely interesting split, but I'd push back a little on how it's framed as a multi-teacher 'merge' — functionally this is just summed task vectors in logit space, and Ilharco's task arithmetic paper out of Washington, 2023, showed you get similar composability just averaging weight-space deltas. 25 00:10:36,050 --> 00:11:01,575 [Hal Turing] I don't think that's the same thing at all, though. Ilharco composes static weight deltas offline, once, and you're stuck with whatever direction you baked in. This is happening per-token, at inference time inside the OPD loop, conditioned on the actual rollout — the student is being corrected live on states it visits, not just handed an averaged checkpoint. That's a different mechanism, not just a different space to do arithmetic in. 26 00:11:01,575 --> 00:11:23,000 [Dr. Ada Shannon] Fair, the online per-token correction against self-generated rollouts is real and task arithmetic doesn't have that. I'll grant the mechanism differs — I still think the compositionality itself, the fact that independent directions just add without much interference, is the less novel half of the story here. Anyway — that mechanism gap is exactly the kind of thing worth pressing on, along with where this whole approach might actually break. 27 00:11:23,000 --> 00:12:04,649 [Hal Turing] Okay, so given all that, let's poke at this a bit, because there's a mechanical detail that bugged me the whole time I was reading. The proxy teacher is literally built by adding and subtracting raw logits across three models — base, positive, negative. That only works if all three share the exact same tokenizer and vocabulary. And every single experiment in this paper is Qwen3-on-Qwen3 — 8B student, 4B, 4B-RL, 0.6B sources, all one family. So when the intro talks about 'abundant existing weak models' like there's this huge open pool to draw from, is that actually true, or is it quietly restricted to same-family checkpoints the whole time? 28 00:12:04,649 --> 00:12:47,375 [Dr. Ada Shannon] It's restricted, and the paper never says so out loud. Logit-space subtraction requires identical vocab indices lining up token-for-token — you can't subtract a Llama logit from a Mistral logit, the index for a given token string isn't even the same number. So this structurally can't touch closed-API models, since you don't get raw logits from a hosted endpoint at all, and it can't touch open models from a different family without some kind of vocabulary alignment step nobody's proposed. Compare that to trajectory-level or black-box weak-to-strong methods, which just need text in and text out. 'Abundant weak models' quietly means 'abundant weak models from your own model family that you can run locally with logit access.' Much narrower claim than the framing suggests. 29 00:12:47,375 --> 00:13:18,725 [Hal Turing] Right, and that same-family framing bleeds into something else — the headline claim, 'the student surpasses the domain teacher it learns from.' That's true for math, 51.8 versus the 4B-RL teacher's 48.8. But look at Table 2 for code: student lands at 60.9, the teacher's sitting at 61.5. The student doesn't surpass the teacher there, it just gets close. And the abstract states the surpassing claim generally, not 'surpassing on math specifically.' 30 00:13:18,725 --> 00:13:39,024 [Dr. Ada Shannon] I'd actually push back a little on how big a deal that is, Hal. It's directionally true — W2S-OPD beats OPD on both domains, and it narrows the code gap substantially even without crossing it. I read the abstract as making the strong claim primarily off the math result and being a bit loose with 'the domain teacher' as a category. Sloppy, sure, but not dishonest. 31 00:13:39,024 --> 00:14:05,799 [Hal Turing] I actually disagree with you there, Ada. It's not just loose wording — it's the load-bearing sentence of the whole paper. If I only read the abstract, I walk away thinking the student beats its teacher, full stop, on both domains they tested. That's the kind of claim that gets cited by other papers without anyone checking Table 2. Generalizing your best result into your headline framing is exactly the overselling Rule 6 tells us to call out. 32 00:14:05,799 --> 00:14:22,349 [Dr. Ada Shannon] Okay — fair, and put that way I'll concede the point. It's not fabrication, but it is the abstract doing the generalizing the data doesn't support, and that's worse than an isolated result caveat buried in a table, because it's the sentence everyone will remember. 33 00:14:22,349 --> 00:14:48,699 [Hal Turing] Sorry to cut you off, but there's a bigger scale question hiding behind that one too. Every experiment here uses an 8-billion-parameter student. The paper's entire motivation is the frontier — models where no larger teacher exists at all — but 8B isn't frontier, it's mid-size. Nobody shows this working when the student is 70B or larger, or when the gap between student and available weak sources is much wider than 4B-to-8B. 34 00:14:48,699 --> 00:15:26,224 [Dr. Ada Shannon] And that's not a nitpick, because we actually have prior work predicting this could break. The Distillation Scaling Laws paper, Busbridge, Ramapuram, Ablin and colleagues out of Apple, 2024, shows the compute-optimal teacher-to-student size ratio shifts as you scale up — a fixed small teacher's signal doesn't stay equally informative as the student grows. Applied here, a capability direction extracted from a 0.6-to-4B contrast pair might just get diluted or misaligned once you're trying to steer a 70B-plus student. Nothing in this paper tests that, and it's exactly the regime the motivation is written for. 35 00:15:26,224 --> 00:16:06,599 [Hal Turing] There's also a citation gap I noticed on the mechanism itself. The paper traces its decoding-time lineage through DExperts and proxy-tuning, both from Liu and colleagues. But the closest actual precedent for 'subtract an amateur's logits from an expert's logits to isolate a direction' is Contrastive Decoding — Li, Holtzman, Fried, and a long list of coauthors including Zettlemoyer and Lewis, 2023. That paper does expert-minus-amateur subtraction at decoding time for text quality, not capability transfer, but it's the same core operation this whole method is built on, and it's just missing from the related work. 36 00:16:06,599 --> 00:16:49,149 [Dr. Ada Shannon] Good catch. There's one more blind spot worth naming before we wrap the critique: those exact small Qwen3 checkpoints — 0.6B, 4B — are already doing a job in production, as speculative-decoding drafters, the small model that proposes tokens a big model verifies. That's the Leviathan, Kalman, and Matias line of work from 2023. The paper never asks whether a model already serving as your drafter can or should double as a contrast source, or whether harvesting a capability direction from it has any implications for its other role. Practically, though, the upside is real — if you've already got small checkpoints and a cheap RL run at small scale lying around, you get a training signal for your big model for nearly free, no need to wait on a bigger teacher or fund a full MOPD-style expert buildout. 37 00:16:49,149 --> 00:17:26,499 [Hal Turing] Which points to where this should go next — can the tokenizer constraint be relaxed by moving from logit-space contrasts to representation-space contrasts, so you could actually pull from a different model family and get closer to the 'abundant weak models' promise in the abstract? And separately, nobody yet knows where weak-to-strong supervision saturates — how many rounds of this you can stack before the signal runs dry, or whether it holds up at genuinely frontier scale, which is the open question the Apple scaling-law work would actually help answer if someone ran the numbers. 38 00:17:26,499 --> 00:17:48,649 [Dr. Ada Shannon] Those are the right open questions. The core idea holds up, though — treating weak-to-strong learning as isolating a transferable direction rather than imitating a weak model directly is a genuinely useful reframing, and it works across three quite different contrast recipes. Just don't read the abstract as 'this generalizes broadly across model families and scales,' because the evidence is same-family, mid-size, and domain-specific. 39 00:17:48,649 --> 00:18:10,299 [Hal Turing] Good summary. So: a clever mechanism, real gains on math, softer on code, untested at frontier scale, and quietly narrower in scope than it sounds. That's Weak-to-Strong On-Policy Distillation from Fangxu Yu and colleagues at Maryland, Microsoft Research, and MBZUAI. Thanks for listening, everyone — we'll catch you next time.