1 00:00:01,000 --> 00:00:43,575 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called 'Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation,' from Yuanyi Wang plus eight co-authors — nine in total — out of The Hong Kong Polytechnic University, InfiX.ai, and the Daya Bay Technology and Innovation Research Institute. It went up on arXiv May 26th, 2026. Ada, here's the number that grabbed me before we even get to the mechanics: they claim you can throw out 95% of your training tokens, keep just the top 5%, and still match or beat training on the full set. 2 00:00:43,575 --> 00:01:16,849 [Dr. Ada Shannon] It shouldn't work, and that tension is why it's worth taking seriously instead of filing it under 'another pruning trick.' The setup is on-policy distillation: a small model trains on text it generated itself, while a bigger teacher grades every token along the way. Recent selective methods assume the wasted part is redundancy — find where teacher and student disagree most, train on that, skip the rest. This paper's claim is sharper than 'better heuristic.' It's that disagreement and learnability are different things entirely, and every prior method scoring tokens by raw disagreement has been measuring the wrong quantity. 3 00:01:16,849 --> 00:01:48,875 [Hal Turing] Let's ground this for anyone new to distillation, since it matters for what's coming. Classic knowledge distillation is the Hinton, Vinyals, and Dean paper out of Google, 2015 — train a small student to match a big teacher's softened output probabilities instead of the one-hot correct label, because the full distribution carries extra information, 'dark knowledge.' That's been the default for a decade. But this paper is specifically about on-policy distillation. What actually separates that from the version most people learned first? 4 00:01:48,875 --> 00:02:30,650 [Dr. Ada Shannon] The textbook version trains the student on text the teacher or dataset wrote — grading itself against someone else's answer key, on someone else's questions. The failure mode is exposure bias: at inference time the student generates its own tokens and keeps going conditioned on its own mistakes, and if it's never practiced recovering from those, small slips compound over a long generation. On-policy distillation fixes that by having the student generate its own rollouts, with the teacher grading every token along that self-generated path using its full next-token distribution. The Agarwal and Vieillard line out of Google DeepMind, 2024 — generalized knowledge distillation — really established this. Dense correction, but on the student's actual trajectory. 5 00:02:30,650 --> 00:02:50,950 [Hal Turing] So the teacher's grading every token the student writes, live, at every position. That already sounds expensive — you need the teacher queryable, or its log-probabilities cached, for every rollout. And the metric doing the grading is KL divergence. How does that number actually capture 'these two models disagree here'? 6 00:02:50,950 --> 00:03:25,325 [Dr. Ada Shannon] That's the tradeoff — OPD needs a live teacher, so it's pricier than static-dataset distillation, but cheaper and lower-variance than full RL, because instead of one scalar reward at the end of a rollout, you get a dense correction at every position. Kullback-Leibler divergence is the tool: given the same context, it measures how far apart the teacher's and student's predicted next-token distributions are. Large KL, big disagreement. Small KL, they basically agree. That number's been the default proxy the whole field uses to decide which tokens deserve supervision — until this paper starts poking at whether it's actually the right one. 7 00:03:25,325 --> 00:03:45,300 [Hal Turing] Oh wait, wait, wait — hold on, that's exactly where this is going, isn't it? Big KL doesn't automatically mean the correction is something the student can actually use. That feels obvious in hindsight, but apparently nobody had rigorously shown it before, they'd just kept using raw disagreement as the stand-in. 8 00:03:45,300 --> 00:04:25,824 [Dr. Ada Shannon] Exactly, and Figure 1 makes it concrete. Picture a token where the student's torn between a handful of candidates — its local support. If the teacher's correction just reweights those same candidates, that's learnable disagreement: a small, actionable nudge toward something the student was already entertaining. But you can get the exact same raw KL value from a token where the teacher's mass lands almost entirely outside anything the student was considering. Same number on paper, completely different situation underneath — the paper calls that incompatible disagreement. You can't meaningfully nudge a model toward an option it currently assigns near-zero probability. Raw KL can't tell those cases apart; it only reports a magnitude. That's the whole reframing. 9 00:04:25,824 --> 00:04:48,099 [Hal Turing] So that's the term — token teachability. Not how big the disagreement looks, but whether the student can actually act on it. And they don't stop at diagnosis — they build a method around it, TA-OPD, Teachability-Aware OPD, that trains only on the positions where that teachability signal runs high, instead of wherever raw KL happens to spike. 10 00:04:48,099 --> 00:05:21,499 [Dr. Ada Shannon] And the headline result is the number you opened with — competitive with, often beating, full-token OPD using only about 5% of the tokens. We won't get into the scoring formula or the diagnostic protocol behind it today — that's its own conversation — but the reframing is the interesting part right now. Selective distillation has spent its short life chasing salience: entropy, raw disagreement, whatever looks biggest. What this paper argues is that the thing that actually matters is whether the student can do anything with the correction it's handed. Salient and learnable turn out to be very different properties. 11 00:05:21,499 --> 00:05:45,374 [Hal Turing] Right, so before TA-OPD even gets built, they have to prove teachability is real and not just a story fit to the data after the fact. That's where this fixed-context diagnostic comes in, and I think it's the cleverest bit of experimental design in the paper. Ada, walk me through what they actually do — because just watching end-to-end benchmark scores go up doesn't tell you which tokens caused it. 12 00:05:45,374 --> 00:06:37,874 [Dr. Ada Shannon] Right, so they freeze a bank of student-generated prefixes before any training happens, then they rescore the exact same prefixes against the teacher twice — once with the original student weights, once after a training update. Same context both times, so any change in the KL gap between student and teacher is attributable to that one gradient step, not to the student wandering into a different rollout. Then, at each frozen position, they look at the student's top-K candidates and the teacher's top-K candidates, take the union, and measure disagreement inside that shared support — that's Dt. Separately, they measure how much teacher probability mass actually lands inside the student's own top-K — that's Ct, compatibility. Multiply the normalized versions together and you get learnable disagreement, DL. The leftover, high Dt but low Ct, is incompatible disagreement, DI — the teacher's correction the student literally can't reach from where it's standing. 13 00:06:37,874 --> 00:07:03,824 [Hal Turing] So a token can look identical on paper — same raw KL, same eye-catching entropy profile — and land in completely different buckets depending on whether the teacher's preferred answer is even on the student's radar. That maps onto something prior work already flagged, right? TIP's Q3 region, low entropy but high KL, the confident-but-wrong zone that selective methods have been chasing as their prime real estate. 14 00:07:03,824 --> 00:07:35,374 [Dr. Ada Shannon] Exactly, and this is where the paper's actual bite is. They show Q3 is not one thing — it's a mixture. Run the same regression inside just that quadrant and you find high-DL tokens inside Q3 produce real fixed-context gain, while low-DL tokens in that same quadrant are flat or even negative, and high-DI tokens are actively wasted budget. Prior selectors were treating the whole quadrant as uniformly golden supervision because it satisfied a two-axis rule, entropy plus divergence. But roughly — 15 00:07:35,374 --> 00:07:49,474 [Hal Turing] Wait, hold on — so TIP's whole selling point, grab the confident-disagreement tokens, is scooping up a bucket that's secretly half wasted signal and calling it a win because the good half drags the average up? 16 00:07:49,474 --> 00:08:30,900 [Dr. Ada Shannon] That's the claim, yes. Which is exactly why they build the teachability score instead of trusting the quadrant boundary. It's steach, normalized Dt times normalized Ct — both put on a zero-to-one scale per rollout batch using robust quantile clipping so one outlier prompt doesn't blow up the ranking. TA-OPD just keeps the top-n positions by that score and masks the OPD loss everywhere else. Then they stress-test it against budget-matched rivals: full OPD with every token, entropy-only, TIP's combined entropy-plus-divergence score, plain random sampling, the incompatible-D score as a deliberately bad control, and a TA-OPD-plus-entropy hybrid to see if mixing in uncertainty helps or just dilutes the signal. 17 00:08:30,900 --> 00:08:50,100 [Hal Turing] And the headline number, the 10% token budget beating full OPD across four teacher-student pairs — how'd that actually break down pair by pair? Because 'competitive with full OPD' is the kind of phrase that can hide a lot of variance if you don't look at where the wins are concentrated. 18 00:08:50,100 --> 00:09:44,050 [Dr. Ada Shannon] Table 3 has TA-OPD winning the average in all four settings — Qwen3-4B to 1.7B, the GRPO-tuned Qwen3-8B to 4B, Qwen3-14B to 4B, and the cross-backbone DeepSeek-R1-Distill-14B to Qwen2.5-3B. But the size of the win isn't flat. The smaller or more mismatched the pair, the bigger the gap over full OPD — the 4B-to-1.7B setting jumps from 42.37 to 44.89, and the cross-backbone pair is the only one where full OPD actually loses to the untouched base model, while TA-OPD stays positive. When the pairing is easy — a strong same-family teacher into a capable student — full OPD and TA-OPD converge, because there's less incompatible junk in the supervision to filter out in the first place. 19 00:09:44,050 --> 00:10:01,100 [Hal Turing] So the messier the teacher-student relationship, the more the filtering actually earns its keep. That naturally raises the budget question — is this a case where more retained tokens just keeps compounding the advantage, or does it plateau? Did they actually sweep it? 20 00:10:01,100 --> 00:10:59,700 [Dr. Ada Shannon] Table 4, yes — 5, 10, 30, and 50 percent, on the two Qwen3-4B student settings. And it's not monotonic at all. For the 8B-GRPO teacher, TA-OPD-plus-entropy actually peaks at 5%, TIP takes 30%, and by 50% TA-OPD alone is back on top. For the 14B teacher, the best average shows up at 10% and then degrades at 30 and 50 — 30% actually dips below the 10% number before creeping back up a hair at 50%. So you've got two Qwen3-4B student settings, same paper, same authors, and the two budget curves don't even agree on direction. TIP wins outright at 30% for the GRPO-tuned 8B teacher, while for the 14B teacher TA-OPD's own edge shrinks or flips the moment you leave the 10% neighborhood. That's not the tidy 'quality beats quantity' story the abstract wants to tell — it's 'the best selector and the best budget are both specific to this teacher-student pair,' which is a much weaker claim. 21 00:10:59,700 --> 00:11:36,425 [Hal Turing] Which makes me want to look harder at how confident we should even be in the headline numbers themselves. Take the 14B-to-4B row in Table 3: TA-OPD averages 54.65, TIP averages 53.62. That's a one-point gap. But the per-benchmark standard deviations on the individual scores run anywhere from about 1.5 points up to 5, and that's over five seeds. So is a one-point gap in the average actually telling us TA-OPD is better, or is it sitting comfortably inside the noise band of a five-seed run? 22 00:11:36,425 --> 00:12:14,900 [Dr. Ada Shannon] Nowhere in the paper do they run a significance test on the averages — no paired bootstrap across seeds, no confidence interval on the gap itself, just point estimates with per-benchmark std shown separately. And once you eyeball it that way, a lot of the 'wins' in Table 3 are single points separating methods whose individual benchmark scores already overlap within one standard deviation of each other. It doesn't mean the effect isn't real — the fixed-context diagnostic in Table 1 does report proper bootstrap intervals that clear zero — but the downstream benchmark table, which is the one that actually matters to a practitioner deciding what to deploy, doesn't get the same statistical rigor. 23 00:12:14,900 --> 00:12:51,225 [Hal Turing] Wait, hold on — that actually connects to something that bugged me about the method itself. To compute Dt and Ct for every position, you need the teacher's top-K log-probs at that position, before you ever apply the 5% mask. So the teacher still has to run a full forward pass over a hundred percent of the tokens to figure out which five percent are worth keeping. Doesn't that mean the '5% retained tokens' framing is only cutting the student's backward pass, while the teacher-side inference cost — which is the expensive half of OPD — doesn't shrink at all? 24 00:12:51,225 --> 00:13:37,225 [Dr. Ada Shannon] That's exactly right, and to their credit the paper says so explicitly in the Limitations section — token budgets refer to KL-supervised positions, not wall-clock savings, and the reported ratio is a supervision budget, not a proportional speedup claim. But that caveat is one sentence at the very end, while the abstract leads with 'often matches full-token OPD with only 5% retained tokens,' which reads like an efficiency result to anyone skimming it. The generalization story has the same shape. Every teacher-student pair is Qwen-family, every training prompt comes from DAPO math data, and the student never exceeds 4B parameters. The one cross-backbone test uses DeepSeek-R1-Distill-Qwen-14B, Daya Guo and colleagues' distillation of the DeepSeek-R1 reasoning model from early 2025, as the teacher — which is one data point, not a generality proof. 25 00:13:37,225 --> 00:14:27,075 [Hal Turing] And even the core mechanism has softer edges than the framing suggests. Their own regression says learnable disagreement's coefficient is only about twice incompatible disagreement's — 0.086 versus 0.043 — and both are positive. So 'incompatible' tokens still carry some signal; it's a dial, not an on-off switch, even though the paper's language treats learnable-versus-incompatible as a clean split. Then in Table 3, HumanEval and IFEval bounce around unpredictably — sometimes TA-OPD beats pure OPD, sometimes it's a wash or worse — on a model trained entirely on math prompts. Ada, is that teachability doing something, or just math-heavy OPD generically perturbing whatever the base model does on code and instruction-following? 26 00:14:27,075 --> 00:14:58,925 [Dr. Ada Shannon] Probably some of both, and the paper doesn't isolate it. There's also the Q3+high-C row buried in Table 2 — the simple heuristic from TIP, Yuanda Xu and colleagues, 2026, of just taking the confident-disagreement quadrant and filtering for high compatibility. It gets close to TA-OPD's fixed-context gain in some settings without needing the full continuous teachability score. So part of the paper's real contribution might just be 'stop over-targeting the region TIP over-targets,' dressed up in a more elaborate formalism than it strictly needs. 27 00:14:58,925 --> 00:15:59,425 [Hal Turing] There's a bigger missed opportunity too, though — this Dt-and-Ct decomposition is really a general per-token calibration diagnostic: does the teacher's correction land somewhere the student can actually reach? That's useful for hallucination detection, reward shaping in RLHF, curriculum ordering, all sorts of places people currently just throw raw entropy or KL at. Shenzhi Wang and colleagues found something structurally similar with high-entropy minority tokens driving RL for reasoning, 2026, but in the RLVR setting rather than distillation. Nobody here tests whether teachability transfers to any of that. And there's a harder wall: the method needs white-box access to the teacher's top-K distribution at every position, so it's built for open teachers only — it says nothing about distilling from a closed, API-only frontier model where you get sampled tokens, not full log-probs. 28 00:15:59,425 --> 00:16:34,475 [Dr. Ada Shannon] So where does that leave practitioners? If you're already running OPD with a queryable teacher — open-weight, on your own cluster — and your bottleneck is the student's training compute or memory for backprop, TA-OPD is a reasonable lever, especially when the student is small or badly mismatched with the teacher, which is exactly where Table 3's gains are largest. But don't reach for it expecting to cut your teacher-inference bill, and don't assume the win generalizes past Qwen-family math distillation until someone runs it on dialogue, multilingual data, or a genuinely different backbone family. 29 00:16:34,475 --> 00:17:12,901 [Hal Turing] That's a fair line to leave it on. The idea itself — that not all teacher-student disagreement is equally absorbable, and you should measure absorbability rather than just salience — is a genuinely useful reframing, and the fixed-context diagnostic is a clean way to isolate it from rollout noise. But the downstream numbers backing it up are thinner than the abstract implies, the budget sweep undercuts the 'quality over quantity' slogan, and the efficiency story only holds for half the compute. Worth reading for the diagnostic, worth being skeptical of for the specific 5% headline. Thanks for listening, and we'll catch you next time.