This episode examines "Learning to Solve Hard Problems in RL for LLMs by Never Giving Up," which finds that RL post-training on math, code, and agentic coding benchmarks disproportionately improves performance on problems models already handle reasonably well, while barely moving the needle on the hardest tasks — a pattern the authors dub the Matthew Effect, after Robert Merton's sociology of science concept. The discussion contrasts two explanations for this skew: the intuitive "signal loss" account, where GRPO's group-relative reward gives zero gradient when every sampled completion fails a hard problem, versus the paper's "signal efficiency" argument, that compute is instead wasted reconfirming easy problems the model has already solved. That distinction matters because it points to different fixes — simply sampling more completions per prompt doesn't help, but reallocating sampling toward unsolved problems does, which motivates the paper's proposed method of persistently re-sampling unsolved prompts rather than discarding them. Listeners get a concrete walkthrough of the controlled GSM8k experiment used to test these competing hypotheses, including a difficulty-tiered evaluation and a K-sweep that challenges conventional assumptions about RL sample scaling. The episode is a useful listen for anyone weighing whether reinforcement learning can actually push language models past their pretrained capability ceiling, or whether it's mainly sharpening skills the model already has.
Sources:
1. Learning to Solve Hard Problems in RL for LLMs by Never Giving Up — Michael Noukhovitch, Hamish Ivison, Nathan Lambert, Aaron Courville, 2026
http://arxiv.org/abs/2609.13443v12. Prioritized Experience Replay — Tom Schaul, John Quan, Ioannis Antonoglou, David Silver, 2015
https://scholar.google.com/scholar?q=Prioritized+Experience+Replay3. DAPO: An Open-Source LLM Reinforcement Learning System at Scale — Qiying Yu, Zheng Zhang, et al. (ByteDance Seed), 2025
https://scholar.google.com/scholar?q=DAPO%3A+An+Open-Source+LLM+Reinforcement+Learning+System+at+Scale4. Prioritized Level Replay — Minqi Jiang, Edward Grefenstette, Tim Rocktäschel, 2021
https://scholar.google.com/scholar?q=Prioritized+Level+Replay5. Curriculum Learning — Yoshua Bengio, Jérôme Louradour, Ronan Collobert, Jason Weston, 2009
https://scholar.google.com/scholar?q=Curriculum+Learning6. Never Give Up: Learning Directed Exploration Strategies — A. P. Badia, P. Sprechmann, A. Vitvitskyi, D. Guo, B. Piot, S. Kapturowski, O. Tieleman, M. Arjovsky, A. Pritzel, A. Bolt, C. Blundell, 2020
https://scholar.google.com/scholar?q=Never+Give+Up%3A+Learning+Directed+Exploration+Strategies7. Reinforce-Ada: An Adaptive Sampling Framework under Non-linear RL Objectives — W. Xiong, C. Ye, B. Liao, H. Dong, X. Xu, C. Monz, J. Bian, N. Jiang, T. Zhang, 2025
https://scholar.google.com/scholar?q=Reinforce-Ada%3A+An+Adaptive+Sampling+Framework+under+Non-linear+RL+Objectives8. POLARIS: A Post-Training Recipe for Scaling Reinforcement Learning on Advanced Reasoning Models — C. An, Z. Xie, X. Li, L. Li, J. Zhang, S. Gong, M. Zhong, J. Xu, X. Qiu, M. Wang, L. Kong, 2025
https://scholar.google.com/scholar?q=POLARIS%3A+A+Post-Training+Recipe+for+Scaling+Reinforcement+Learning+on+Advanced+Reasoning+Models9. RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs? — Y. Sun, Y. Cao, P. Huang, H. Bai, H. Hajishirzi, N. Dziri, D. Song, 2025
https://scholar.google.com/scholar?q=RL+Grokking+Recipe%3A+How+Does+RL+Unlock+and+Transfer+New+Algorithms+in+LLMs%3F10. The Art of Scaling Reinforcement Learning Compute for LLMs — D. Khatri, L. Madaan, R. Tiwari, R. Bansal, S. S. Duvvuri, M. Zaheer, I. S. Dhillon, D. Brandfonbrener, R. Agarwal, 2025
https://scholar.google.com/scholar?q=The+Art+of+Scaling+Reinforcement+Learning+Compute+for+LLMs11. The Primacy Bias in Deep Reinforcement Learning — E. Nikishin, M. Schwarzer, P. D'Oro, P.-L. Bacon, A. Courville, 2022
https://scholar.google.com/scholar?q=The+Primacy+Bias+in+Deep+Reinforcement+Learning12. Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts (GRESO) — H. Zheng, Y. Zhou, B. R. Bartoldson, B. Kailkhura, F. Lai, J. Zhao, B. Chen, 2025
https://scholar.google.com/scholar?q=Act+Only+When+It+Pays%3A+Efficient+Reinforcement+Learning+for+LLM+Reasoning+via+Selective+Rollouts+%28GRESO%29Interactive Visualization: The Matthew Effect in RL: Rich Get Richer on Hard Problems