AI Post Transformers · Episode Companion

Dream-RSI: Teaching AI How to Search, Not Just Solve

Tong Zheng, Xidong Wu, Zheng Zhang, Zhankui He, Chaoyi Zhang, Benjamin Coleman, et al. (17 authors) — Google, University of Maryland College Park, Google DeepMind, University of Virginia · posted September 14, 2026
arXiv:2609.14858 Recursive Self-Improvement Exploration Policy Optimization Discovery Trees

What gets optimized here

Classic recursive self-improvement (RSI) refines the candidate — the code, proof, or design. Dream-RSI keeps the underlying coding agent (Gemini) frozen and instead recursively improves the exploration policy: the orchestration layer deciding which branches to expand, how to batch workers, and when to stop.

Online Explore — real Gemini calls, tree grows Construct Simulator — freeze tree as replay Dream — score M candidate policies for free

Fixed vs. Learned Exploration

Bandit heuristics (UCB, epsilon-greedy) fix the exploration rule once. RL² (2016) and Never Give Up (2020) first showed exploration itself could be learned. Dream-RSI applies that to LLM-driven code/algorithm discovery, where FunSearch (2024) is the fixed-strategy baseline it moves past.

Why "Dreaming" Is Cheap

A finished discovery run leaves behind a tree: every generate-evaluate attempt, its parent, artifact, and score. Replaying that tree to test a new policy costs nothing — no new agent calls, just re-reading recorded outcomes.

0
agent calls per dream
M
candidate policies / round
≥v0
best never regresses

Replaying a Discovery Tree

Each round the policy selects up to W nodes (leaves or root) for the real agent to expand during online exploration. During dreaming, the same decision interface is reused, but instead of calling Gemini, the system hands back whatever child was already recorded — deterministic, zero cost. A policy can only walk paths that already exist; it can't invent a branch nobody explored.

Revealed / recorded node Newly touched this round Unexplored in tree

Scoring a Dreamed Policy — Three Terms

What ranks the M candidate policies inside a dream:

Lasso Regularization Path — Agent Calls to Match/Beat Baselines

Dream-RSI beats sklearn and glmnet on all six held-out datasets, matching or beating SimpleTES's runtime (51,200 generations) using roughly two orders of magnitude fewer calls. The edge over its own fixed-exploration baseline is real but smaller.

Math Optimization — Table 1 Heatmap

Hover a cell. Sum-Difference is a fourth-decimal gap. Autocorrelation is a SimpleTES win. Circle Packing is a six-way tie to six decimals — Dream-RSI gets there with <1,000 generations vs. SimpleTES's 51,200.

KernelBench — Efficiency Multipliers

Matched-performance generation savings (VGG16, LayerNorm) and matched-budget score gains (ConvDiv, ConvMax) over Recursive Fixed Exploration.

ConvDiv: Policy Adapting Mid-Run

Round-by-round attempts evaluated on ConvDiv: the policy cuts effort as performance climbs, then ramps back up once progress plateaus.

What the Headline Numbers Leave Out

Every reported call count (317, 1879, "<1,000 generations") counts only calls to the coding agent during online exploration. It excludes the policy-development agent, which makes its own LLM call for every one of the M candidate policy revisions produced during dreaming, every round.

Unaccounted: replay itself is free (reading stored nodes), but producing each candidate policy version requires an LLM call to read replay feedback and rewrite code. That cost sits entirely outside the reported 162× SimpleTES speedup and the 317-vs-550 edge over fixed exploration — and the paper never reports it.

Single-Seed, Single-Vendor

No error bars, no repeated trials anywhere in the paper. All eight tasks ran exclusively on Gemini-3.1 Pro / Gemini-3.7-Flash via the Gemini CLI — entirely inside Google/DeepMind. Whether the learned policy-writing skill transfers to Claude or GPT as the underlying discovery agent is untested.

Fixed vs. Learned Policy Lineage

AlphaEvolve (2025) is the fixed-policy predecessor. EvoX (2026) optimizes exploration policy online — exactly the delayed-feedback problem Dream-RSI routes around via offline replay. Dreamer V3's world model imagines unobserved states; Dream-RSI's "world" only resequences recorded history.

References

1Dream-RSI: Recursive Self-Improvement through Evolving Worlds — Zheng, Wu, Zhang, et al., 2026
2RL²: Fast Reinforcement Learning via Slow Reinforcement Learning — Duan, Schulman, Chen, Bartlett, Sutskever, Abbeel, 2016
3Never Give Up: Learning Directed Exploration Strategies — Badia, Sprechmann, Vitvitskyi, et al. (DeepMind), 2020
4Mathematical discoveries from program search with LLMs (FunSearch) — Romera-Paredes, Barekatain, Novikov, et al. (DeepMind), 2024
5Taking the Human Out of the Loop: A Review of Bayesian Optimization — Shahriari, Swersky, Wang, Adams, de Freitas, 2016
6AlphaEvolve: A coding agent for scientific and algorithmic discovery — Novikov et al. (DeepMind), 2025
7Mastering diverse domains through world models (Dreamer V3) — Hafner, Pasukonis, Ba, Lillicrap, 2023
8EvoX: Meta-evolution for automated discovery — Liu et al., 2026
9Evaluation-driven scaling for scientific discovery (SimpleTES) — Ye et al., 2026