Classic recursive self-improvement (RSI) refines the candidate — the code, proof, or design. Dream-RSI keeps the underlying coding agent (Gemini) frozen and instead recursively improves the exploration policy: the orchestration layer deciding which branches to expand, how to batch workers, and when to stop.
Bandit heuristics (UCB, epsilon-greedy) fix the exploration rule once. RL² (2016) and Never Give Up (2020) first showed exploration itself could be learned. Dream-RSI applies that to LLM-driven code/algorithm discovery, where FunSearch (2024) is the fixed-strategy baseline it moves past.
A finished discovery run leaves behind a tree: every generate-evaluate attempt, its parent, artifact, and score. Replaying that tree to test a new policy costs nothing — no new agent calls, just re-reading recorded outcomes.
Each round the policy selects up to W nodes (leaves or root) for the real agent to expand during online exploration. During dreaming, the same decision interface is reused, but instead of calling Gemini, the system hands back whatever child was already recorded — deterministic, zero cost. A policy can only walk paths that already exist; it can't invent a branch nobody explored.
What ranks the M candidate policies inside a dream:
Dream-RSI beats sklearn and glmnet on all six held-out datasets, matching or beating SimpleTES's runtime (51,200 generations) using roughly two orders of magnitude fewer calls. The edge over its own fixed-exploration baseline is real but smaller.
Hover a cell. Sum-Difference is a fourth-decimal gap. Autocorrelation is a SimpleTES win. Circle Packing is a six-way tie to six decimals — Dream-RSI gets there with <1,000 generations vs. SimpleTES's 51,200.
Matched-performance generation savings (VGG16, LayerNorm) and matched-budget score gains (ConvDiv, ConvMax) over Recursive Fixed Exploration.
Round-by-round attempts evaluated on ConvDiv: the policy cuts effort as performance climbs, then ramps back up once progress plateaus.
Every reported call count (317, 1879, "<1,000 generations") counts only calls to the coding agent during online exploration. It excludes the policy-development agent, which makes its own LLM call for every one of the M candidate policy revisions produced during dreaming, every round.
No error bars, no repeated trials anywhere in the paper. All eight tasks ran exclusively on Gemini-3.1 Pro / Gemini-3.7-Flash via the Gemini CLI — entirely inside Google/DeepMind. Whether the learned policy-writing skill transfers to Claude or GPT as the underlying discovery agent is untested.
AlphaEvolve (2025) is the fixed-policy predecessor. EvoX (2026) optimizes exploration policy online — exactly the delayed-feedback problem Dream-RSI routes around via offline replay. Dreamer V3's world model imagines unobserved states; Dream-RSI's "world" only resequences recorded history.
| 1 | Dream-RSI: Recursive Self-Improvement through Evolving Worlds — Zheng, Wu, Zhang, et al., 2026 | arXiv:2609.14858 |
| 2 | RL²: Fast Reinforcement Learning via Slow Reinforcement Learning — Duan, Schulman, Chen, Bartlett, Sutskever, Abbeel, 2016 | Scholar → |
| 3 | Never Give Up: Learning Directed Exploration Strategies — Badia, Sprechmann, Vitvitskyi, et al. (DeepMind), 2020 | Scholar → |
| 4 | Mathematical discoveries from program search with LLMs (FunSearch) — Romera-Paredes, Barekatain, Novikov, et al. (DeepMind), 2024 | Scholar → |
| 5 | Taking the Human Out of the Loop: A Review of Bayesian Optimization — Shahriari, Swersky, Wang, Adams, de Freitas, 2016 | Scholar → |
| 6 | AlphaEvolve: A coding agent for scientific and algorithmic discovery — Novikov et al. (DeepMind), 2025 | Scholar → |
| 7 | Mastering diverse domains through world models (Dreamer V3) — Hafner, Pasukonis, Ba, Lillicrap, 2023 | Scholar → |
| 8 | EvoX: Meta-evolution for automated discovery — Liu et al., 2026 | Scholar → |
| 9 | Evaluation-driven scaling for scientific discovery (SimpleTES) — Ye et al., 2026 | Scholar → |