This episode explores "The Unlearnability Phenomenon in RLVR for Language Models" by Yulin Chen and colleagues at NYU, which uncovers a puzzling failure mode in reinforcement learning with verifiable reward (RLVR)—the training method underlying reasoning models like o1, o3, DeepSeek-R1, and QwQ. The hosts unpack how GRPO, the algorithm popularized by DeepSeek, relies on reward variance across sampled rollouts to compute learning signals, and how the paper's authors tracked individual hard training examples to discover that some receive genuine positive reward repeatedly yet never show improved success rates—even after training converges. The discussion probes why this defies basic policy-gradient intuition, since a rewarded rollout should become more probable regardless of whether the model got the right answer through skill or luck. The core investigative thread centers on gradient cosine similarity—checking whether an example's own learning signal aligns with or fights against the rest of the training batch—as the lens for explaining why some correctly-solved problems never stick. Listeners interested in the mechanics and hidden limits of frontier reasoning-model training will find this a sharp look at a ceiling effect invisible in ordinary loss curves.
Sources:
1. The Unlearnability Phenomenon in RLVR for Language Models — Yulin Chen, He He, Chen Zhao, 2026
http://arxiv.org/abs/2605.167872. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models — Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Mingchuan Zhang, Y.K. Li, Y. Wu, Daya Guo (DeepSeek-AI), 2024
https://scholar.google.com/scholar?q=DeepSeekMath%3A+Pushing+the+Limits+of+Mathematical+Reasoning+in+Open+Language+Models3. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — DeepSeek-AI (Daya Guo et al.), 2025
https://scholar.google.com/scholar?q=DeepSeek-R1%3A+Incentivizing+Reasoning+Capability+in+LLMs+via+Reinforcement+Learning4. DAPO: An Open-Source LLM Reinforcement Learning System at Scale — Qiying Yu, Zheng Zhang, Yu Yue, Mingxuan Wang, et al. (ByteDance Seed / Tsinghua AIR), 2025
https://scholar.google.com/scholar?q=DAPO%3A+An+Open-Source+LLM+Reinforcement+Learning+System+at+Scale5. Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? — Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, et al. (Tsinghua University), 2025
https://scholar.google.com/scholar?q=Does+Reinforcement+Learning+Really+Incentivize+Reasoning+Capacity+in+LLMs+Beyond+the+Base+Model%3F6. OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling — Z. Wang, F. Zhou, X. Li, P. Liu, 2025
https://scholar.google.com/scholar?q=OctoThinker%3A+Mid-training+Incentivizes+Reinforcement+Learning+Scaling7. Arithmetic without Algorithms: Language Models Solve Math with a Bag of Heuristics — Y. Nikankin, A. Reusch, A. Mueller, Y. Belinkov, 2025 (ICLR)
https://scholar.google.com/scholar?q=Arithmetic+without+Algorithms%3A+Language+Models+Solve+Math+with+a+Bag+of+Heuristics8. Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification — A. Zhang, Y. Chen, J. Pan, C. Zhao, A. Panda, J. Li, H. He, 2025
https://scholar.google.com/scholar?q=Reasoning+Models+Know+When+They%27re+Right%3A+Probing+Hidden+States+for+Self-Verification9. The Invisible Leash: Why RLVR May or May Not Escape Its Origin — F. Wu, W. Xuan, X. Lu, M. Liu, Y. Dong, Z. Harchaoui, Y. Choi, 2026
https://scholar.google.com/scholar?q=The+Invisible+Leash%3A+Why+RLVR+May+or+May+Not+Escape+Its+Origin10. The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models — G. Cui, Y. Zhang, J. Chen, L. Yuan, Z. Wang, Y. Zuo, et al., 2025
https://scholar.google.com/scholar?q=The+Entropy+Mechanism+of+Reinforcement+Learning+for+Reasoning+Language+ModelsInteractive Visualization: The Unlearnability Phenomenon in RLVR Reasoning Models