This episode examines "Dream-RSI: Recursive Self-Improvement through Evolving Worlds," which proposes making the exploration strategy of a discovery system — not the underlying model — the target of recursive self-improvement. Rather than optimizing candidate solutions directly, the system optimizes the policy that decides how to search: which branches to expand, how to parallelize workers, and when to stop, while the underlying coding agent (Gemini, in the paper's experiments) stays fixed. The key innovation is "dreaming": replaying an already-recorded discovery tree of past generate-evaluate attempts as a cheap simulator, letting new exploration policies be scored for free against historical outcomes instead of running costly new agent calls. This produces a three-stage loop — online exploration to grow the tree, constructing a replay simulator from it, then "dreaming" to test and select better policies before redeploying them — addressing the classic problem that policy-level exploration research suffers from painfully delayed feedback. The discussion situates the work against prior exploration methods like bandit algorithms, RL², Never Give Up, and FunSearch, making it a useful listen for anyone interested in how search-strategy meta-optimization, rather than raw model capability, might be the next lever for scaling AI-driven discovery.
Sources:
1. Dream-RSI: Recursive Self-Improvement through Evolving Worlds — Tong Zheng, Xidong Wu, Zheng Zhang, Zhankui He, Chaoyi Zhang, Benjamin Coleman, Ruoqiao Wei, Di Bai, Haolin Liu, Rui Liu, Xue Wang, Yue Zhuan, Wang-Cheng Kang, Renkai Xiang, Heng Huang, Xinwu Cheng, Yunsong Guo, 2026
http://arxiv.org/abs/2609.148582. RL²: Fast Reinforcement Learning via Slow Reinforcement Learning — Yan Duan, John Schulman, Xi Chen, Peter L. Bartlett, Ilya Sutskever, Pieter Abbeel, 2016
https://scholar.google.com/scholar?q=RL%C2%B2%3A+Fast+Reinforcement+Learning+via+Slow+Reinforcement+Learning3. Never Give Up: Learning Directed Exploration Strategies — Adrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, et al. (DeepMind), 2020
https://scholar.google.com/scholar?q=Never+Give+Up%3A+Learning+Directed+Exploration+Strategies4. Mathematical discoveries from program search with large language models (FunSearch) — Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, et al. (Google DeepMind), 2024
https://scholar.google.com/scholar?q=Mathematical+discoveries+from+program+search+with+large+language+models+%28FunSearch%295. Taking the Human Out of the Loop: A Review of Bayesian Optimization — Bobak Shahriari, Kevin Swersky, Ziyu Wang, Ryan P. Adams, Nando de Freitas, 2016
https://scholar.google.com/scholar?q=Taking+the+Human+Out+of+the+Loop%3A+A+Review+of+Bayesian+Optimization6. AlphaEvolve: A coding agent for scientific and algorithmic discovery — A. Novikov et al., 2025
https://scholar.google.com/scholar?q=AlphaEvolve%3A+A+coding+agent+for+scientific+and+algorithmic+discovery7. Mastering diverse domains through world models (Dreamer V3) — D. Hafner, J. Pasukonis, J. Ba, T. Lillicrap, 2023
https://scholar.google.com/scholar?q=Mastering+diverse+domains+through+world+models+%28Dreamer+V3%298. EvoX: Meta-evolution for automated discovery — S. Liu et al., 2026
https://scholar.google.com/scholar?q=EvoX%3A+Meta-evolution+for+automated+discovery9. Evaluation-driven scaling for scientific discovery (SimpleTES) — H. Ye et al., 2026
https://scholar.google.com/scholar?q=Evaluation-driven+scaling+for+scientific+discovery+%28SimpleTES%29Interactive Visualization: Dream-RSI: Teaching AI How to Search, Not Just Solve