RLVR / GRPO training loop
Every "reasoning" model this year — o1, o3, DeepSeek-R1, QwQ — is post-trained with this loop. Hover a stage for detail. The orange branch is the failure mode dynamic sampling is built to route around — and it's not the phenomenon this paper is about.
GRPO (Group Relative Policy Optimization, DeepSeekMath 2024) replaces PPO's learned critic with a simple trick: sample a group of rollouts per prompt, standardize the reward within that group. The entire learning signal lives in the spread within a group — never the absolute pass rate.
How big is the unlearnable slice?
"Unlearnable" = received at least one positive reward during training, yet scores below 10% pass@1 (32 samples, <3 successes) after convergence. It solved the problem — it just never learned from solving it.
Three pairings, zero overlap in setting: Qwen2.5-0.5B on MATH-Easy, Llama-3.2-3B-Instruct on MATH-Hard, Qwen2.5-3B on DeepScaleR (40k problems). The curve view shows the strange part — reward keeps landing on the unlearnable line, and the success rate just won't move.
Gradient cosine similarity — the actual diagnostic
For each correct rollout, compute its GRPO gradient and measure cosine similarity against the rest of the batch. Hover a cell for the exact reading.
Easy examples cluster tightly at ≈0.75 with each other. Unlearnable examples don't cluster with anything — not even a 3×3 diagonal shows them agreeing with themselves, sitting at ≈0.46. Each one is effectively its own island in gradient space, so every isolated update gets diluted or undone by the next batch.
Every obvious fix, tried and killed
Rollout scarcity, clipping/KL suppression, and semantic-duplicate augmentation — all applied to the unlearnable group specifically. Toggle the metric; both stay flat.
Oversampling + replay guarantees a correct rollout every batch — no movement. Clip-higher with the KL term removed — no movement. GPT-5-generated near-duplicate problems — pass@1 ticks up briefly then plateaus, pass@16 actually drops (overfitting), and gradient similarity to the near-identical twin stays low. Semantic closeness doesn't buy optimization closeness.
The one lever that moved it: mid-training
Llama-3.2-3B-Base vs. two OctoThinker-3B variants — same base checkpoint lineage, only variable is 20B tokens of mid-training with a different data mix. No RL touches these models at this point.
Gradient similarity on hard MATH examples is substantially higher for both mid-trained variants than for the untouched base. This is the fault line in the paper's own framing: RL-stage fixes failing tells you GRPO can't repair a gap it didn't create — it doesn't prove the gap survives at frontier scale, or that it's a "fundamental limitation of RL post-training" rather than a pretraining-exposure problem.