This episode explores the RAGEN-2 paper’s claim that agentic reinforcement learning can produce reasoning traces that look active and diverse while losing real dependence on the input. It explains the paper’s central distinction between ordinary entropy, which measures diversity within a single prompt, and template collapse, where traces across many different prompts become generic variations of the same pattern. The discussion also covers the proposed mutual-information-style monitoring approach, which rescoring traces against other prompts to test whether reasoning remains identifiable to its source, and links the failure mode to weak reward signal, PPO-style regularization, and sparse long-horizon feedback. Listeners would find it interesting because it reframes a core question in reasoning RL: not whether an agent looks busy, but whether its reasoning is still actually about the problem in front of it.
Sources:
1. RAGEN-2: Reasoning Collapse in Agentic RL — Zihan Wang, Chi Gui, Xing Jin, Qineng Wang, Licheng Liu, Kangrui Wang, Shiqi Chen, Linjie Li, Zhengyuan Yang, Pingyue Zhang, Yiping Lu, Jiajun Wu, Li Fei-Fei, Lijuan Wang, Yejin Choi, Manling Li, 2026
http://arxiv.org/abs/2604.062682. MINE: Mutual Information Neural Estimation — Mohamed Ishmael Belghazi, Aristide Baratin, Sai Rajeswar, Sherjil Ozair, Yoshua Bengio, Aaron Courville, R. Devon Hjelm, 2018
https://scholar.google.com/scholar?q=MINE%3A+Mutual+Information+Neural+Estimation3. Representation Learning with Contrastive Predictive Coding — Aaron van den Oord, Yazhe Li, Oriol Vinyals, 2018
https://scholar.google.com/scholar?q=Representation+Learning+with+Contrastive+Predictive+Coding4. Learning deep representations by mutual information estimation and maximization — R. Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Phil Bachman, Adam Trischler, Yoshua Bengio, 2019
https://scholar.google.com/scholar?q=Learning+deep+representations+by+mutual+information+estimation+and+maximization5. RAGEN-2: Reasoning Collapse in Agentic RL — Zihan Wang, Chi Gui, Xing Jin, Qineng Wang, Manling Li, et al., 2026
https://scholar.google.com/scholar?q=RAGEN-2%3A+Reasoning+Collapse+in+Agentic+RL6. Proximal Policy Optimization Algorithms — John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, Oleg Klimov, 2017
https://scholar.google.com/scholar?q=Proximal+Policy+Optimization+Algorithms7. Signal-to-Noise Ratio Analysis of Policy Gradient Algorithms — John W. Roberts, Russ Tedrake, 2008
https://scholar.google.com/scholar?q=Signal-to-Noise+Ratio+Analysis+of+Policy+Gradient+Algorithms8. Understanding the Impact of Entropy on Policy Optimization — Zafarali Ahmed, Nicolas Le Roux, Mohammad Norouzi, Dale Schuurmans, 2019
https://scholar.google.com/scholar?q=Understanding+the+Impact+of+Entropy+on+Policy+Optimization9. Prioritized Experience Replay — Tom Schaul, John Quan, Ioannis Antonoglou, David Silver, 2016
https://scholar.google.com/scholar?q=Prioritized+Experience+Replay10. DAPO: An Open-Source LLM Reinforcement Learning System at Scale — Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, et al., 2025
https://scholar.google.com/scholar?q=DAPO%3A+An+Open-Source+LLM+Reinforcement+Learning+System+at+Scale11. No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping — Thanh-Long V. Le, Myeongho Jeon, Kim Vu, Viet Lai, Eunho Yang, 2025
https://scholar.google.com/scholar?q=No+Prompt+Left+Behind%3A+Exploiting+Zero-Variance+Prompts+in+LLM+Reinforcement+Learning+via+Entropy-Guided+Advantage+Shaping12. RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning — Zihan Wang et al., 2025
https://scholar.google.com/scholar?q=RAGEN%3A+Understanding+Self-Evolution+in+LLM+Agents+via+Multi-Turn+Reinforcement+Learning13. The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models — Ganqu Cui et al., 2025
https://scholar.google.com/scholar?q=The+Entropy+Mechanism+of+Reinforcement+Learning+for+Reasoning+Language+Models14. ASTER: Agentic Scaling with Tool-integrated Extended Reasoning — Xuqin Zhang, Quan He, Zhenrui Zheng, Zongzhang Zhang, Xu He, and Dong Li, 2026
https://scholar.google.com/scholar?q=ASTER%3A+Agentic+Scaling+with+Tool-integrated+Extended+Reasoning15. Demystifying Reinforcement Learning in Agentic Reasoning — Zhaochen Yu, Ling Yang, Jiaru Zou, Shuicheng Yan, and Mengdi Wang, 2025
https://scholar.google.com/scholar?q=Demystifying+Reinforcement+Learning+in+Agentic+Reasoning16. The Price of Format: Diversity Collapse in LLMs — Longfei Yun, Chenyang An, Zilong Wang, Letian Peng, and Jingbo Shang, 2025
https://scholar.google.com/scholar?q=The+Price+of+Format%3A+Diversity+Collapse+in+LLMs17. Revisiting Entropy in Reinforcement Learning for Large Reasoning Models — Renren Jin et al., 2025
https://scholar.google.com/scholar?q=Revisiting+Entropy+in+Reinforcement+Learning+for+Large+Reasoning+Models18. ERPO: Token-Level Entropy-Regulated Policy Optimization for Large Reasoning Models — Song Yu and Li Li, 2026
https://scholar.google.com/scholar?q=ERPO%3A+Token-Level+Entropy-Regulated+Policy+Optimization+for+Large+Reasoning+Models19. Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning — Chen Qian et al., 2025
https://scholar.google.com/scholar?q=Demystifying+Reasoning+Dynamics+with+Mutual+Information%3A+Thinking+Tokens+are+Information+Peaks+in+LLM+Reasoning20. MITS: Enhanced Tree Search Reasoning for LLMs via Pointwise Mutual Information — Jiaxi Li et al., 2025
https://scholar.google.com/scholar?q=MITS%3A+Enhanced+Tree+Search+Reasoning+for+LLMs+via+Pointwise+Mutual+Information21. On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning — Yifan Zhang et al., 2025
https://scholar.google.com/scholar?q=On+the+Design+of+KL-Regularized+Policy+Gradient+Algorithms+for+LLM+Reasoning22. Rethinking KL Regularization in RLHF: From Value Estimation to Gradient Optimization — Kezhao Liu et al., 2025
https://scholar.google.com/scholar?q=Rethinking+KL+Regularization+in+RLHF%3A+From+Value+Estimation+to+Gradient+Optimization23. What Makes a Reward Model a Good Teacher? An Optimization Perspective — Noam Razin et al., 2025
https://scholar.google.com/scholar?q=What+Makes+a+Reward+Model+a+Good+Teacher%3F+An+Optimization+Perspective24. Accelerating RLHF Training with Reward Variance Increase — Zonglin Yang et al., 2025
https://scholar.google.com/scholar?q=Accelerating+RLHF+Training+with+Reward+Variance+Increase25. Efficient RLVR Training via Weighted Mutual Information Data Selection — Xinyu Zhou et al., 2026
https://scholar.google.com/scholar?q=Efficient+RLVR+Training+via+Weighted+Mutual+Information+Data+Selection26. AI Post Transformers: Stabilizing Efficient Reasoning with Step-Level Advantage Selection — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-29-stabilizing-efficient-reasoning-with-ste-1e589d.mp327. AI Post Transformers: Reasoning Theater and Unfaithful Chain-of-Thought — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-reasoning-theater-and-unfaithful-chain-o-a4507e.mp328. AI Post Transformers: Benchmarking Test-Time Scaling for General LLM Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-benchmarking-test-time-scaling-for-gener-8f14f9.mp329. AI Post Transformers: Can Models Learn from Long Context? — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-can-models-learn-from-long-context-77533e.mp3