1 00:00:01,000 --> 00:00:41,850 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into "Dream-RSI: Recursive Self-Improvement through Evolving Worlds," by Tong Zheng et al. — seventeen authors total — out of Google, University of Maryland College Park, Google DeepMind, and University of Virginia, posted to arXiv on September 14th, 2026. And Ada, the thing that jumped out at me isn't the result, it's the framing: they're not trying to make the AI better at solving problems. They're trying to make the AI better at deciding *how to search* for solutions in the first place. 2 00:00:41,850 --> 00:01:11,724 [Dr. Ada Shannon] Right, and that's a weirder target than it sounds. Most of the recursive self-improvement work we've seen treats the search process as fixed plumbing — you set your strategy once, and then you just pour more compute through it. This paper says: what if the strategy itself is the thing you recursively improve? So instead of one loop — generate, evaluate, refine the solution — you've got a loop wrapped around a loop. The inner loop is still generating candidate algorithms or kernels. The outer loop is asking, 'is the way I'm exploring even any good, and can I make it better without burning a fortune finding out?' 3 00:01:11,724 --> 00:01:22,049 [Hal Turing] Okay, let's actually define that term for people, because 'recursive self-improvement' gets thrown around a lot. What do they mean by it here, precisely? 4 00:01:22,049 --> 00:01:57,150 [Dr. Ada Shannon] Classic RSI is: an agent generates candidates, evaluates outcomes, incorporates feedback, and refines future iterations — that loop applied to whatever the agent is producing, code, proofs, designs. Here they take that exact loop and apply it one level up, to the exploration policy — the orchestration layer that decides which branches of a search to expand, how many workers run in parallel, when to stop. Normally that layer is hand-tuned once and frozen. Dream-RSI makes it programmable and lets the system rewrite it based on its own track record, while the actual coding agent underneath — in their experiments, Gemini — never changes. 5 00:01:57,150 --> 00:02:03,700 [Hal Turing] So it's not the model getting smarter, it's the search strategy around the model getting smarter. 6 00:02:03,700 --> 00:02:42,525 [Dr. Ada Shannon] Exactly. And that distinction matters because of where the real bottleneck sits. They point to two failure modes in prior systems. One: fixed exploration strategies can't adapt as the search space scales — you pick a heuristic up front and it's stuck being right or wrong for the rest of the run. Two: if you try to optimize the exploration policy online, while discovery is happening, you're stuck with brutally delayed feedback. You don't know if a new exploration strategy was good until you've run it for potentially thousands of proposal-evaluation cycles. That's expensive and slow, and the space of possible policies is huge, so trial and error at that level is a terrible way to spend compute. 7 00:02:42,525 --> 00:03:00,325 [Hal Turing] Which is where I actually want to push back a little going in — that sounds like the classic exploration-exploitation problem dressed up in agent language. Epsilon-greedy, UCB bandits, that whole toolbox has been handling 'how do I explore' for decades. What's actually new here? 8 00:03:00,325 --> 00:03:57,275 [Dr. Ada Shannon] Fair pushback, and it connects to something worth explaining properly, because this podcast hasn't covered it before. What you're describing — epsilon-greedy, upper-confidence-bound bonuses, Thompson sampling — those are fixed heuristics. Somebody picks the exploration rule once, tunes a constant, and it runs unchanged. Exploration Policy Optimization is the idea of treating that rule itself as a learnable object. It's not new in concept — Duan and Schulman's RL-squared paper out of Berkeley and OpenAI back in 2016 showed you could train a policy that learns to explore, and DeepMind's Never Give Up work in 2020 mixed multiple exploration strategies via a bandit inside Agent57. What's different here is applying that to LLM-driven code and algorithm discovery specifically, where FunSearch, the DeepMind paper from Romera-Paredes and colleagues in 2024, is the fixed-strategy baseline this line of work is trying to move past. 9 00:03:57,275 --> 00:04:09,425 [Hal Turing] Oh wait wait wait — so their actual claim isn't 'we invented a better search heuristic,' it's 'we figured out how to get the meta-optimization loop cheap enough to be worth running at all.' 10 00:04:09,425 --> 00:05:02,150 [Dr. Ada Shannon] That's the whole paper in one sentence, yeah. And their trick for getting it cheap is almost sneaky in its simplicity. Once a discovery run finishes, you're sitting on a discovery tree — a root plus every generate-evaluate attempt branching off it, each node storing its parent, the artifact it produced, and its score. Their insight is that this tree is already a usable simulator. You can replay it — try a different order of visiting branches, different parallel groupings, different stopping points — and because every outcome was already recorded, you get instant feedback without calling the actual coding agent again. They call this 'dreaming,' loosely borrowing the framing from Ha and Schmidhuber's World Models and Hafner's Dreamer line of work, where an agent imagines rollouts inside a learned model instead of the real environment. It's a looser analogy than it sounds — more on that later — but the intuition carries: explore inside a cheap replica before you spend real money exploring the real thing. 11 00:05:02,150 --> 00:05:09,875 [Hal Turing] It's basically studying a map of a city you already walked through, instead of walking it again to test a new route. 12 00:05:09,875 --> 00:05:31,375 [Dr. Ada Shannon] That's the right picture. And that gives you the three-stage loop the whole system runs: online explore, where the current policy drives the real coding agent and grows the discovery tree; construct the replay simulator from that tree; then dream — evaluate a bunch of candidate policies for free against that replay, pick the best one, and redeploy it for the next round of real exploration. Each cycle, the simulator pool grows, so the dreaming gets richer too. 13 00:05:31,375 --> 00:05:39,750 [Hal Turing] Alright — so next up, let's get into how that dreaming step actually works mechanically, and whether the numbers back up the pitch. 14 00:05:39,750 --> 00:06:21,375 [Dr. Ada Shannon] So mechanically, online rollout is simple: each round the policy picks a batch of up to W nodes — leaves or the root — and hands them to the real Gemini coding agent, which resumes each node's saved workspace and produces one new attempt per node. The evaluator scores it, the attempts become new children, the tree grows. That's the only place real agent calls happen. Dreaming reuses that same decision interface, but instead of generating anything new, it replays — pick a batch, and the system hands back whatever child was already recorded there. Deterministic, zero agent calls. They run M candidate policy versions this way against the accumulated trees, and a separate policy-development agent rewrites the policy's code between versions based on how each scored. 15 00:06:21,375 --> 00:06:35,875 [Hal Turing] Okay so the constraint is: whatever the new policy does, it can only walk paths that already exist in the tree — it can't invent a branch nobody explored. What's actually being scored as 'good' in that replay? 16 00:06:35,875 --> 00:07:07,550 [Dr. Ada Shannon] The best version always beats or ties the one it replaced — that's guaranteed, because the current policy is quietly included as version zero in every batch, so the max can't go down. What ranks them is a three-part score: the best quality found anywhere in the revealed subtree, a penalty scaled by how many nodes got revealed, since burning through the tree isn't free even in a dream, and a bonus for revealing many nodes per decision round — rewarding batching attempts in parallel instead of working through them one at a time. 17 00:07:07,550 --> 00:07:20,350 [Hal Turing] And they test this across three pretty different domains — Lasso solvers, math optimization, and GPU kernels. Start with Lasso, since that's where the efficiency numbers are loudest. 18 00:07:20,350 --> 00:08:12,500 [Dr. Ada Shannon] Right, the Lasso regularization path task, benchmarked against sklearn — Pedregosa and colleagues, out of INRIA, 2011 — and glmnet, Friedman, Hastie and Tibshirani, Stanford, 2010 — plus SimpleTES, Ye and colleagues, 2026, which burns 51,200 generations. Dream-RSI's discovered solvers beat both standard libraries on all six held-out datasets, and match or beat SimpleTES's runtime using between 317 and 1879 discovery-agent calls depending on the backbone — roughly two orders of magnitude fewer than SimpleTES. It also edges past Recursive Fixed Exploration on call count: 317 versus 550 on Gemini-3.1 Pro, 1879 versus 3200 on Gemini-3.7-Flash. 19 00:08:12,500 --> 00:08:32,625 [Hal Turing] Oh wait wait wait — hold on, two orders of magnitude against SimpleTES is the headline, sure, but 317 versus 550 against their own fixed-exploration baseline is less than half, not two orders of magnitude. Does that margin hold up in the math benchmarks too? 20 00:08:32,625 --> 00:09:25,125 [Dr. Ada Shannon] Barely, and it's worth sitting with. Table 1 has Dream-RSI at 1.145427 on Sum-Difference against Recursive Fixed Exploration's 1.144047 — a fourth-decimal-place gap. On Autocorrelation, where lower is better, SimpleTES actually takes the outright win at 1.453675 against Dream-RSI's 1.456375, but it needed 51,200 generations to get there. Circle Packing is the strangest: Dream-RSI, Recursive Fixed Exploration, SimpleTES, and three more baselines all land at 2.635983 — a six-way tie to six decimal places — with Dream-RSI using under a thousand generations instead of SimpleTES's 51,200. 21 00:09:25,125 --> 00:09:34,575 [Hal Turing] So on the math side it's less 'we won' and more 'we tied at a fraction of the cost.' What about kernel engineering — does the story hold up there? 22 00:09:34,575 --> 00:10:16,300 [Dr. Ada Shannon] That one's cleaner. On KernelBench's VGG16 and LayerNorm, Dream-RSI matches final performance with 2.43 and 1.79 times fewer generations than Recursive Fixed Exploration. On ConvDiv and ConvMax, at matched budget, it scores 2.09 and 1.44 times higher. There's also a good ablation: compressing history into a text guidance summary instead of an interactive replay simulator consistently underperformed replay — heavy semantic steering over-constrains the search. And there's a nice trace on ConvDiv of the policy adapting round to round: it cuts evaluated attempts from 110 to 50 while performance climbs, then ramps effort back up once progress plateaus. 23 00:10:16,300 --> 00:10:41,550 [Hal Turing] Okay, here's my problem with the accounting, Ada. Every number we just cited — 317 calls, 1879 calls, under a thousand generations — only counts calls to the coding agent during online exploration. It doesn't count what the policy-development agent is doing during dreaming, rewriting the exploration policy's code across M revisions, every single round. 24 00:10:41,550 --> 00:11:03,375 [Dr. Ada Shannon] Exactly, and the paper never reports that cost. M is a free parameter, dreaming happens every outer round, and each revision is another call to the development agent reading trajectories and rewriting policy code. Across five or ten rounds with a nontrivial M, that's real compute sitting outside the headline number — and it's genuinely unclear whether the edge over SimpleTES or fixed exploration survives once it's counted. 25 00:11:03,375 --> 00:11:22,675 [Hal Turing] And that compounds with the other thing — this is all single-seed, one run per method per task. On an LLM-driven search process, a gap that small between 1.145427 and 1.144047 could just be noise from a different sampling seed. 26 00:11:22,675 --> 00:11:42,350 [Dr. Ada Shannon] Right, no error bars, no repeated trials anywhere in the paper. Which is honestly why their own framing — 'matches or surpasses' — is more defensible than treating this as a clean win. The efficiency story on Lasso and the kernels looks real. The math-optimization story looks more like a tie dressed up as a victory. 27 00:11:42,350 --> 00:12:17,524 [Hal Turing] And that accounting gap is what actually worries me more than the tie, Ada — so the thing doing the real intellectual labor here — deciding how to rewrite the policy after reading the replay feedback — is a separate LLM making its own string of calls, and none of that shows up in the ledger at all? That's like a company reporting factory output but leaving the R&D department's salaries off the balance sheet. If dreaming is supposed to be the cheap half of this loop, I'd want to see exactly how cheap once you count the agent doing the dreaming, not just the agent doing the building. 28 00:12:17,524 --> 00:12:50,600 [Dr. Ada Shannon] Right, and to be precise, replay itself is free — it's just reading stored nodes. But producing each of the M candidate policy versions requires an LLM call to read replay feedback and rewrite code. That's real inference sitting entirely outside the headline metric. Add it back in, and does the 162x over SimpleTES hold? Does the 317-versus-550 edge over Recursive Fixed Exploration survive? We don't know, and in a paper whose entire pitch is 'we're dramatically cheaper,' that's not a footnote — it's the thing that determines whether the pitch is even true. 29 00:12:50,600 --> 00:13:26,774 [Hal Turing] Hold on, wait — there's a second thing tangled up in that same efficiency pitch that I don't think we can let slide, which is that every one of these eight tasks ran exclusively on Gemini-3.1 Pro or Gemini-3.7-Flash, through the Gemini CLI, entirely inside Google and DeepMind. Doesn't that make it genuinely hard to know how much of this is 'a good general exploration policy' versus 'a policy that happens to exploit quirks in how Gemini specifically handles agentic tool calls and long context windows'? 30 00:13:26,774 --> 00:14:05,924 [Dr. Ada Shannon] That's exactly the generalization question I'd want answered before recommending anyone adopt this. The orchestration layer itself is model-agnostic — it just hands batches of nodes to whatever coding agent you plug in underneath. But the policy-development agent is an LLM learning to write better exploration code by reading Gemini-shaped replay traces. Whether that skill transfers cleanly to Claude or GPT as the underlying discovery agent, without retuning the whole loop, is completely untested. It's a single-vendor validation of a framework that's pitched, in the abstract, as general infrastructure for any agentic discovery setup. 31 00:14:05,924 --> 00:14:29,099 [Hal Turing] Zooming out to where this sits historically — you've mentioned AlphaEvolve a couple times already as the thing their math baselines are actually measured against. Walk me through that contrast properly, and then I want to hear how loose that Dreamer comparison in their own framing really is, because 'dreaming' is doing a lot of marketing work sitting right there in the title of this paper. 32 00:14:29,099 --> 00:15:49,699 [Dr. Ada Shannon] AlphaEvolve — Novikov and colleagues out of Google DeepMind, 2025 — is the fixed-exploration-policy predecessor: a strong evolutionary coding agent that never questions its own search strategy. Several AlphaEvolve-family variants land within noise of Dream-RSI on Circle Packing in that same Table 1, so it's really fixed versus learned exploration policy converging on the same ceiling. On Dreamer — that's Dreamer V3, Hafner, Pasukonis, Ba, and Lillicrap, 2023 — the analogy is looser than it sounds. Dreamer learns a generative dynamics model that can imagine states it's never observed. Dream-RSI's 'world' is a literal replay of discrete outcomes already recorded; it can resequence history, never invent new history. And against EvoX, Liu and colleagues, 2026 — the most direct prior work treating exploration strategy itself as an optimization target — the real distinction is offline replay versus EvoX's online policy optimization, which is exactly the delayed-feedback problem Dream-RSI is trying to route around. Practically, that's still the appeal here regardless of the accounting gaps: it bolts on top of a frozen coding agent as a control layer, no fine-tuning required, so the barrier to trying it is genuinely low for a team already generating discovery trees worth replaying. 33 00:15:49,699 --> 00:16:23,249 [Hal Turing] Fair summary to close on: the core idea — treat a completed discovery history as a free simulator instead of throwing it away — is genuinely clever and worth stealing regardless of how the specific numbers shake out. But the missing meta-optimization compute in their cost accounting, and the single-seed, single-vendor results in Table 1, mean the size of the real advantage here is still an open question rather than a settled one. Ada, thanks for digging through the tables with me on this one. That's it for today, everyone — thanks for listening, and we'll catch you next time.