1 00:00:01,000 --> 00:00:49,125 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper is 'The Unlearnability Phenomenon in RLVR for Language Models,' by Yulin Chen et al. — three authors total, with He He and Chen Zhao — out of New York University and NYU Shanghai. It hit arXiv May 16th this year, headed to ICML 2026 in Seoul. Here's the line that stopped me in the abstract: a substantial subset of hard training examples stays permanently unlearnable — even when the model gets them right, repeatedly, during training. Not 'never solves it.' Solves it, over and over, and the success rate just doesn't move. Ada, in reinforcement learning, doesn't getting rewarded for something usually mean you do more of it? 2 00:00:49,125 --> 00:01:19,299 [Dr. Ada Shannon] That's exactly the assumption this paper detonates. RLVR is the training recipe underneath basically every 'reasoning' model you've used this year — OpenAI's o1 and o3, DeepSeek's R1, Qwen's QwQ, Kimi. If a real chunk of your hard training examples are structurally immune to learning, that's not a curiosity, that's a ceiling on how far RL post-training can push reasoning. And it's invisible from the loss curve — the reward signal looks perfectly healthy the whole time while those examples quietly stop improving. 3 00:01:19,299 --> 00:01:37,099 [Hal Turing] Before we go further, let's set the table for anyone who hasn't sat inside an RL post-training run — I don't want to assume knowledge. What is RLVR, mechanically, and why did it basically eat the entire 'how do we build reasoning models' conversation in under two years? 4 00:01:37,099 --> 00:02:21,099 [Dr. Ada Shannon] RLVR — Reinforcement Learning with Verifiable Reward — swaps a human-trained reward model for a cheap automatic checker: does the answer match, does the code pass its tests. No humans scoring outputs like classic RLHF. The model samples several full attempts, 'rollouts,' each gets a reward, and you push it toward whichever rollouts scored well. The algorithm that made this practical at scale is GRPO — Group Relative Policy Optimization — from the DeepSeekMath paper, Shao and coauthors, DeepSeek-AI, 2024. Instead of a separate critic network like PPO uses, GRPO samples a group of rollouts per prompt and standardizes the reward within that group. DeepSeek-R1 — Guo and coauthors, DeepSeek-AI, 2025 — showed it scales to frontier reasoning with zero supervised warm-start. 5 00:02:21,099 --> 00:02:41,625 [Hal Turing] Right, and that z-score trick only works if there's something to standardize. If every rollout in the group is correct, or every single one is wrong, the advantage is zero across the board and the example contributes nothing to that update. So the whole method leans on reward variance existing inside the group. 6 00:02:41,625 --> 00:03:01,275 [Dr. Ada Shannon] Exactly, and that's the assumption everyone builds on — get variance, GRPO computes a real advantage, gradient flows toward the correct rollout, away from the incorrect ones, case closed. Except when Chen and her coauthors actually tracked individual hard examples across training, they found a group that gets that variance, genuinely gets rewarded — 7 00:03:01,275 --> 00:03:15,974 [Hal Turing] Wait — wait, hold on. You're telling me the reward signal is there, GRPO computes a real advantage, the gradient update fires exactly like it's supposed to... and the model still doesn't get better at that specific problem? 8 00:03:15,974 --> 00:03:52,849 [Dr. Ada Shannon] That's the phenomenon. They split hard examples — the ones the model initially struggles with — into two camps. 'Learnable' examples improve smoothly over training. 'Unlearnable' examples consistently receive positive reward during training, yet the success rate never climbs. Their working definition: an example counts as unlearnable if, after training converges, the final policy still scores below ten percent pass-at-one — thirty-two samples, fewer than three successes — despite having received at least one positive reward earlier on. It solved the problem; it just never learned from solving it. 9 00:03:52,849 --> 00:04:11,074 [Hal Turing] Play devil's advocate with me for a second, Ada. If the model stumbles onto a correct answer by luck a couple of times out of who knows how many attempts, of course that doesn't generalize into a stable skill. That's low sample rate, not some deep representational mystery. 10 00:04:11,074 --> 00:04:43,750 [Dr. Ada Shannon] I actually disagree with you there, Hal. Policy gradient doesn't care why a rollout was correct — that's the whole mechanism. You reward whatever the model sampled, and the update nudges the policy to make that behavior more likely next time. Lucky guess or not, a rewarded rollout is supposed to become more probable. So if reward keeps showing up and the success rate refuses to move, either the mechanism is broken for these examples, or there's something about them the model structurally can't absorb. Neither of those is 'just noise.' 11 00:04:43,750 --> 00:05:06,000 [Hal Turing] Fair — the mechanism should reinforce it regardless of whether the reasoning behind it was any good. What's still nagging me is whether that ten-percent threshold is doing some of the work, making borderline examples look permanently stuck. But I'll admit, 'reward keeps showing up and the policy refuses to move toward it' is a genuinely strange sentence. 12 00:05:06,000 --> 00:06:05,100 [Dr. Ada Shannon] That skepticism is worth holding onto — we'll get into how they stress-test that threshold. But here's the thread they end up pulling: instead of trusting the reward numbers alone, they look at the gradients — cosine similarity between the gradient a correct rollout produces and the gradient direction implied by the rest of the training batch. If an example's own gradient points roughly the same way as everything else, learning it and learning the rest of the dataset reinforce each other. If it doesn't, every isolated update gets diluted, or undone, by the next batch. That's the lens they use to answer the question we opened with — why examples with a real, positive reward signal still never seem to stick. The mechanics matter here — GRPO draws k rollouts per prompt, then converts the binary correct-or-incorrect hits into an advantage by subtracting the group's mean reward and dividing by its standard deviation. Credit isn't 'right or wrong,' it's 'right relative to your siblings.' They also run dynamic sampling — if a whole group is right or wrong, the standard deviation is zero, the advantage is undefined, and that prompt gets dropped. 13 00:06:05,100 --> 00:06:24,450 [Hal Turing] So the entire learning signal lives in the spread within a group, never the absolute pass rate — stricter than I expected. Okay, with the mechanism clear, how big is this unlearnable slice once you actually run it for real, across different models and datasets rather than just one narrow setup? 14 00:06:24,450 --> 00:06:56,725 [Dr. Ada Shannon] Bigger than you'd hope. Three pairings — Qwen2.5-0.5B on the easier band of MATH, Llama-3.2-3B-Instruct on the harder band, and Qwen2.5-3B on DeepScaleR, forty thousand problems. Table 1 puts the unlearnable share of hard examples at seventeen to thirty percent, worst for the smallest model at just over thirty. And that's after excluding anything that never got a single positive reward the entire run — this is specifically the 'got rewarded repeatedly, never improved' bucket. 15 00:06:56,725 --> 00:07:17,625 [Hal Turing] Thirty percent of the hard problems, permanently stuck even after they exclude the ones that never got rewarded at all. Walk me through how they actually killed off the obvious explanations here — I'd guess rollout scarcity first, then something about clipping or the KL penalty smothering the signal before it can do anything useful? 16 00:07:17,625 --> 00:07:55,425 [Dr. Ada Shannon] Both, in sequence. First, scarcity — maybe these examples don't get enough correct rollouts per batch. So they force it: oversampling with replay, guaranteeing a fixed one-correct-to-seven-incorrect ratio every batch, replaying buffered correct rollouts when a batch comes up short. So every unlearnable example gets guaranteed positive signal each step — and the gap to the learnable group doesn't close at all. So, suppression next: maybe correct rollouts start from low-probability territory under the reference model and get squashed by clipping or the KL term. They check reference log-likelihood directly, and the distributions overlap almost completely across all three groups. 17 00:07:55,425 --> 00:08:09,225 [Hal Turing] Wait — wait, hold on. So guaranteeing a correct rollout every single batch didn't move the unlearnable group at all, and now the low-probability story is dead too? Both of the tidy explanations just die? 18 00:08:09,225 --> 00:08:45,300 [Dr. Ada Shannon] Both die. They track clipping ratios too — the three groups sit right on top of each other the entire run. Then they retrain with clip-higher and the KL term removed, the standard exploration fixes, and it still doesn't budge the unlearnable group. Two negative results, which sends them into the gradients. For each example they take correct rollouts, compute the GRPO gradient, and measure cosine similarity against the rest of the set. Easy examples cluster tightly, roughly 0.75 similarity within group. Unlearnable examples don't cluster with anything, not even each other — around 0.46 even against their own group. 19 00:08:45,300 --> 00:09:04,450 [Hal Turing] So each unlearnable example is basically its own island in gradient space, not even keeping company with other unlearnable examples — that's a much starker picture than 'somewhat harder to learn.' Did they find anything in the actual reasoning traces that lines up with that isolation? 20 00:09:04,450 --> 00:09:27,125 [Dr. Ada Shannon] They did. They scored reasoning traces for correct-answer rollouts with GPT-5-mini, and unlearnable examples score noticeably worse even with a correct final answer. There's an example — a volume problem in three dimensions — where the model correctly sets up the case analysis, then botches the case enumeration, contradicts its own earlier reasoning, and still lands on the right number. The gap doesn't shrink with training either — by step 120 it's wider than at step 50. 21 00:09:27,125 --> 00:09:46,750 [Hal Turing] Okay, but doesn't that basically mean the model isn't reasoning at all on these examples? It contradicts itself mid-derivation and still gets the right token — that reads like pattern-matching dressed up as reasoning, and the reward signal can't tell the two apart. Isn't that what the paper's basically implying? 22 00:09:46,750 --> 00:10:04,674 [Dr. Ada Shannon] I'd push back on 'isn't reasoning at all,' Hal. It's landing the correct answer well above chance, repeatedly, on problems it's never once solved coherently — that's some kind of narrow, real capability doing the work, not pure noise. Calling it nothing undersells whatever mechanism is producing that answer. 23 00:10:04,674 --> 00:10:19,799 [Hal Turing] Fair — I'll walk 'not reasoning at all' back to 'not reasoning the way the reward assumes.' Either way, outcome-only reward can't tell a genuine derivation from a lucky shortcut. So can they just paper over the gap with more data? 24 00:10:19,799 --> 00:10:48,899 [Dr. Ada Shannon] They tried. GPT-5 generates near-duplicate problems and decomposed subproblems for every unlearnable example, and they train on those alongside the originals. Pass at one ticks up briefly then plateaus early, and pass at sixteen actually drops — classic overfitting on a fixed set. Worse, gradient similarity between an unlearnable example and its near-identical augmented twin stays low. Structurally similar problems, dissimilar gradients — semantic closeness doesn't buy you optimization closeness. 25 00:10:48,899 --> 00:11:02,600 [Hal Turing] So the representation gap survives literally everything they throw at it during RL — clipping, replay, augmentation, all of it. Is there anything that actually moves that number, or is it just a wall? 26 00:11:02,600 --> 00:11:42,500 [Dr. Ada Shannon] One thing does — mid-training. They compare Llama-3.2-3B-Base against two OctoThinker-3B variants that had twenty billion tokens of mid-training on that same base, with different data mixes. Gradient similarity on hard MATH examples comes out substantially higher for both OctoThinker versions than for the untouched base model — something upstream of RL entirely is already reshaping how these examples sit in the model. Same architecture, same base checkpoint lineage, and the only variable is twenty billion tokens of mid-training with a different data mix — no RL touches these models at that point. And that alone moves the exact number they use to diagnose 'flawed representation' in the first place. 27 00:11:42,500 --> 00:12:17,399 [Hal Turing] Hold on, Ada — that's not a footnote, that's a problem for the paper's own framing. If changing only the mid-training diet substantially raises gradient similarity on hard examples, and gradient similarity is literally how they define the representation flaw, doesn't that cut against calling unlearnability a 'fundamental limitation' of RLVR or GRPO? That sounds less like a wall the algorithm hits and more like a symptom of these particular checkpoints not having seen enough of the right pretraining exposure yet. 28 00:12:17,399 --> 00:12:38,875 [Dr. Ada Shannon] I actually disagree with you there, Hal — or at least with 'less like a wall.' It's not either-or. Go back to what they killed: rollout scarcity, clipping and KL suppression, data augmentation with problems that are practically identical in content. Every RL-stage lever failed regardless of what originally caused the flawed representation. That result stands on its own whether or not mid-training would eventually fix it. 29 00:12:38,875 --> 00:13:12,475 [Hal Turing] Sorry to cut in, but — those aren't the same claim, though. 'RL-stage fixes don't work' just tells you GRPO can't repair a gap it didn't create. It says nothing about whether that gap survives at frontier scale. Calling it a 'fundamental limitation of RL post-training for reasoning,' full stop, in the conclusion, is a much stronger claim than 'GRPO alone won't fix a problem inherited from pretraining' — and their own Section 5.4 is evidence for the second, weaker claim, not the first. 30 00:13:12,475 --> 00:13:48,850 [Dr. Ada Shannon] Okay, fair, that's a real distinction and I'll grant you the conclusion overreaches a little. But here's a methodological problem that bugs me more: look at what's actually confounded across their three settings. Qwen2.5-0.5B on MATH Easy, Llama-3.2-3B-Instruct on MATH Hard, Qwen2.5-3B on DeepScaleR — three different model sizes, three completely non-overlapping datasets, no shared setting anywhere. So when they report seventeen to thirty percent unlearnable, there's no way to cleanly say how much is 'small model lacks capacity' versus 'this particular dataset is just harder.' Every knob moved at once. 31 00:13:48,850 --> 00:14:09,675 [Hal Turing] Right, and that same ambiguity bleeds into the reasoning-trace story too. That case-enumeration example you mentioned earlier — a correct answer riding on incoherent intermediate steps — is that really 'no reasoning happening,' or is it a model reaching for a real, separate capability that just doesn't look like textbook derivation? 32 00:14:09,675 --> 00:14:54,775 [Dr. Ada Shannon] That's almost exactly the frame from 'Arithmetic without Algorithms: Language Models Solve Math with a Bag of Heuristics,' Nikankin, Reusch, Mueller, and Belinkov, out of Technion, ICLR 2025 — the paper this one cites for that idea. Their finding is that models solve arithmetic through a genuine grab-bag of learned heuristic circuits, not algorithmic computation. This paper treats that as noise to be eliminated. But there's real tension with another citation in their own Discussion — 'Reasoning Models Know When They're Right,' by Zhang, Chen, Pan, Zhao, Panda, Li, and He, 2025, out of NYU — that's He He and Chen Zhao again, the same two co-advisors on today's paper. That work shows reasoning models carry rich latent self-verification signals internally. So which is it — genuinely flawed representation, or a usable signal the RL objective just isn't reading? 33 00:14:54,775 --> 00:15:03,675 [Hal Turing] So if you're a lab actually running RLVR pipelines right now, what do you do with any of this? Where does the fix actually live? 34 00:15:03,675 --> 00:15:42,825 [Dr. Ada Shannon] Upstream, is the honest answer, given everything that failed here. Clipping, KL, oversampling, curriculum, augmentation — all dead ends at the RL stage, and the one lever that moved the needle was mid-training data composition, before RL ever starts. That's a real, if unglamorous, takeaway: invest in what the base model reads before you optimize how it's rewarded. Credit to the authors, too — their limitations section is upfront that this is small-to-mid-scale models, math only, and that pass-at-one under ten percent is an operational convention, not a hard boundary. I just wish that honesty carried into the abstract's language as much as it does the fine print. 35 00:15:42,825 --> 00:16:23,200 [Hal Turing] That's a good note to land on. Bottom line for anyone catching up: some hard training examples get the right answer, get rewarded, and still never improve — not because of scarce positive signal, not because of clipping or KL, but because the model's internal representation for that example is an outlier that RL alone can't repair. Whether that's truly fundamental to RLVR or mostly a pretraining-exposure problem is still an open question, and the paper's own mid-training result is the strongest evidence pointing at the second answer. Thanks for listening, everyone — we'll catch you next time.