1 00:00:01,000 --> 00:00:45,071 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a survey called "Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops," by Mingguang Chen et al. — three authors total — out of UC Riverside, a startup called AlphaAvatar, and Illinois Institute of Technology, posted to arXiv in July 2026. Ada, here's the number that stopped me before I even got to their arguments: they read and classified twelve hundred and fifty papers. 2 00:00:45,071 --> 00:01:17,161 [Dr. Ada Shannon] That's not a survey, Hal, that's closer to a census. And the detail that matters more than the raw count: seventy-four percent of that corpus was posted in 2026 alone. So this isn't a field looking backward at a settled literature — it's an attempt to map ground that's still shifting under their feet while they're standing on it. They built it in two passes: a seed harvest of eight hundred seventy-one papers across seven threads of self-improvement research, then a targeted supplemental harvest of three hundred seventy-nine more, filling gaps their own taxonomy exposed once they had one. 3 00:01:17,161 --> 00:01:31,557 [Hal Turing] Okay, so before we go further, I need us to actually define terms, because I've seen "self-refine," "self-reward," "self-play," "self-evolve," "self-distill" thrown around like they're interchangeable, and I don't think they are. 4 00:01:31,557 --> 00:02:19,901 [Dr. Ada Shannon] They're not, and that conflation is basically the paper's thesis statement. So let's fix vocabulary. An agent, in their usage, is an LLM in a perceive-act loop — it observes state, takes actions through tools, and iterates until some stopping condition. A single model call is not an agent. The harness, or scaffolding, is everything wrapped around that model — prompts, tools, memory, skill libraries, orchestration code — and critically, it's editable, including by the agent itself. And an evaluator is any mechanism that maps an output to a quality signal. They split that into verifiers, which carry soundness guarantees like a proof checker or a test suite, and judges, which don't — an LLM grading another LLM's homework. 5 00:02:19,901 --> 00:02:23,895 [Hal Turing] That distinction between verifier and judge is going to matter a lot later, isn't it. 6 00:02:23,895 --> 00:03:03,276 [Dr. Ada Shannon] It's the spine of the whole paper. Self-improvement, in their definition, just means a system produces a better version of itself or its outputs where "better" is defined by some evaluator. That dependency on an evaluator isn't a footnote — every failure mode this paper documents traces back to that one clause. Which is why they draw their central cut: bounded self-refinement improves against a fixed, external evaluator, so it's convergent and you can actually measure it. Open-ended recursive self-improvement modifies the system and the criteria for improvement itself, with no fixed anchor — divergent in principle, not just in practice. 7 00:03:03,276 --> 00:03:11,728 [Hal Turing] None of this is new as an idea, though — I mean, people have been worried about a machine bootstrapping itself since before either of us was writing code. 8 00:03:11,728 --> 00:03:50,598 [Dr. Ada Shannon] Since 1965, actually. I. J. Good's "Speculations Concerning the First Ultraintelligent Machine" is the founding document — pure philosophical argument, no system behind it, but it's where "intelligence explosion" comes from: a machine smart enough to improve its own design gets better at improving itself, and that compounds. Then in two thousand three, extended in 2006, Schmidhuber formalized the strong version with his Gödel machines — a self-referential solver that can rewrite any part of itself, including its own rewriting code, but only executes a change once it has an actual mathematical proof that the change improves future performance— 9 00:03:50,598 --> 00:04:01,419 [Hal Turing] Oh wait, hold on — that's the one where the proof requirement is basically what makes it safe AND useless at the same time, right? Provably optimal but you can't generate the proof for real code. 10 00:04:01,419 --> 00:04:50,134 [Dr. Ada Shannon] That's exactly the catch, and it sets up the whole modern era: everything since has been trading Schmidhuber's unattainable proof for something empirically checkable instead — a benchmark score, a passing test — which is weaker but actually runs. The paper explicitly says it's using Anthropic's recent essay on recursive self-improvement as a motivating frame and a source of stage vocabulary, not as evidence — worth flagging since it's an essay, not peer-reviewed data. Anthropic's framing is a continuum from humans writing all code to agents delegating to agents to, eventually, "closing the loop" — agents designing their own successors. And they layer a second axis on top: degree of loop closure. Is a human reviewing every change, auditing an automated signal, or is there no human at all? 11 00:04:50,134 --> 00:04:55,707 [Hal Turing] And underneath all of that sits one more ranking, right — how much you can actually trust the signal doing the judging? 12 00:04:55,707 --> 00:05:38,014 [Dr. Ada Shannon] Exactly — they call it a verification hierarchy: formal verifiers at the top, then execution feedback, then learned judges, then intrinsic self-assessment at the bottom, weakest rung. Two more names worth planting now, we'll come back to them: Lightman et al.'s "Let's Verify Step by Step" established that scoring reasoning steps beats scoring only final answers, and Huang et al.'s 2024 result that LLMs mostly can't self-correct their own reasoning without outside help reshaped everything that followed. Cross those four improvement categories — deployment, training, evaluation, and auto research — against the loop-closure axis, and you get the whole map this paper is about to walk us through. 13 00:05:38,014 --> 00:06:10,429 [Dr. Ada Shannon] —establishes that scoring intermediate reasoning steps beats scoring only the final answer, and that seeded the whole process-reward-model literature. Huang et al.'s 2024 paper is the field's cold shower: absent external feedback, LLMs mostly can't self-correct reasoning, and naive self-correction can make things worse. Hold both of those, the 2026 literature spends two years responding. But first, how this survey actually got built, because the corpus construction here is unusually rigorous for a field this young. 14 00:06:10,429 --> 00:06:19,624 [Hal Turing] Walk me through it, Ada, how do you land on twelve hundred fifty papers without it collapsing into arbitrary keyword soup? That's my worry with any survey at this scale. 15 00:06:19,624 --> 00:07:16,931 [Dr. Ada Shannon] Two stages. Seed harvest: arXiv queries across seven threads, self-refinement, self-rewarding training, automated AI research, self-modifying agents, code and algorithm discovery, RSI theory and safety, self-generated-data loops, enriched with OpenAlex metadata, off-topic bleed stripped, landing at eight hundred seventy-one papers. Then re-classification into the taxonomy using theme defaults plus keyword rules, with manual correction, eighty-nine papers moved off their default category, three more overridden later. Then a targeted supplemental harvest of three hundred seventy-nine papers in directions the seed queries under-covered: self-evaluation, test-time training, zero-data self-play. Twelve fifty total. It breaks down as deployment-time self-evolution at three ninety-three, training-time self-iteration at three forty, self-evaluation at three eighteen, Auto Research at one thirty-nine, foundations at just sixty. 16 00:07:16,931 --> 00:07:28,355 [Hal Turing] Oh wait wait wait, sixty? The category that's supposedly about whether any of this is safe is the smallest bucket by a mile, next to three ninety-three for deployment. File that, I want to come back to it. 17 00:07:28,355 --> 00:08:07,597 [Dr. Ada Shannon] It is, and the paper has more to say about what that skew means later. What I'll flag now is self-evaluation, three eighteen and the fastest-growing slice, eighty-two percent posted in 2026 alone. That's Self-Rewarding Language Models territory, Yuan et al., 2024, out of Meta: the same model generates a response and judges it, policy and reward improving together. Direct ancestor of the process-reward lineage, Lightman's outcome-versus-process split, then ReST-MCTS*, which doesn't filter final answers for correctness, it infers process rewards by running tree search over intermediate steps, so a lucky guess with broken reasoning doesn't get credit. 18 00:08:07,597 --> 00:08:15,817 [Hal Turing] So if improvement strength tracks how trustworthy the judge is, like you set up before the break, what does that look like rung by rung, with actual examples? 19 00:08:15,817 --> 00:09:22,040 [Dr. Ada Shannon] Top rung, formal verifiers, proof checkers, type systems, sound by construction: theorem-proving self-play, verified skill evolution. Next, execution feedback, tests, compilers, reliable but incomplete, since passing tests don't guarantee correctness and any fixed benchmark eventually gets gamed: code self-repair, compiler tuning. Third, learned judges, reward models, LLM-as-judge, bounded by the judge's own competence and themselves an optimization target. Bottom, intrinsic signals, confidence, self-consistency, cheapest and most gameable. FunSearch and AlphaEvolve both sit at the top two rungs. And there's a study measuring the floor directly: the Mirror Loop experiment ran three providers' models through ten rounds of pure ungrounded self-critique across four task families, and informational change declined fifty-five percent, the models were reformulating, not improving. One verification step inserted at round three restored forward movement. 20 00:09:22,040 --> 00:09:34,300 [Hal Turing] Fifty-five percent from nothing but talking to itself. That's Huang's result playing out live. Which brings up the scarier, slower version, what does Shumailov's collapse result actually show? 21 00:09:34,300 --> 00:10:26,220 [Dr. Ada Shannon] Shumailov et al., in Nature, models trained recursively on their own outputs lose the tails of the distribution and degenerate across generations, the statistical version of inbreeding. That's the pessimistic pole. Counter-evidence exists too: diffusion models can train on their own generations once you control perceptual alignment and hallucination accumulation; self-play survives when data gating and reward grounding are managed as separate levers; reasoning self-training collapse traces to fixable data imbalance and overthinking. Two other failure shapes worth naming: the self-confirming loop, where generator and judge share weights so their biases correlate and the model gets over-rewarded for the mistakes it's most confident about; and diversity collapse, where co-evolving proposers converge on the narrow reward band and starve the solver's curriculum. Novelty is a resource these loops burn down. 22 00:10:26,220 --> 00:10:33,093 [Hal Turing] Give me the headline Auto Research results, Ada, the ones where this stops being theoretical and actually starts shipping into production systems. 23 00:10:33,093 --> 00:11:41,035 [Dr. Ada Shannon] AlphaEvolve's the cleanest case, its discoveries fed straight back into Google's own infrastructure: faster matrix-multiplication kernels, data-center scheduling, accelerator circuit simplifications. But A-Evolve-Training is the one that got me, an autonomous system running the entire post-training loop of a thirty-billion-parameter model, proposing data and recipe changes, launching runs, deciding what to keep, four rounds over multiple weeks, zero human in the loop. Mid-run, it noticed its own development metric had decoupled from external performance, candidates driving the internal score up without improving anything real, and revised its own search policy to treat that proxy as evidence against a candidate instead of for one. A system catching its own evaluator's corruption, in the wild. SkillsBench is the reality check on the other side: human-authored skills improve pass rates by sixteen points, LLM-authored skills give no measurable gain. And ScienceAgentBench, ResearchArena, and MLReplicate all converge on the same gap, these systems produce something that looks like research, but feasibility and quality are not the same thing. 24 00:11:41,035 --> 00:12:15,447 [Hal Turing] So that's the anecdote propping up the governance argument, and here's what bugs me about the claim underneath it, Ada. The paper says improvement strength tracks the verification hierarchy, then calls that 'a qualitative pattern we observe throughout the corpus, not a measured law.' That's their own hedge. But sections seven and eight build the entire safety narrative on it anyway. Why not tag all twelve fifty papers by verification tier and correlate that against outcomes, instead of pointing at FunSearch and AlphaEvolve as if two cases settle it? 25 00:12:15,447 --> 00:12:45,725 [Dr. Ada Shannon] There's a theoretical reason the intuition feels solid without that correlation — Gao, Schulman, and Hilton's 2022 reward-overoptimization paper out of OpenAI already gives you the Goodhart curve one level down: push a learned judge hard enough and true quality rises, then falls. So the intuition has a formal backbone. But that's a different claim from testing it across their own corpus, and letting theory stand in for an empirical result they never ran shouldn't anchor a governance argument. 26 00:12:45,725 --> 00:13:14,472 [Hal Turing] There's a second methods problem right under that one. All twelve fifty papers were sorted by a single annotator — theme defaults, keyword rules, manual correction, eighty-nine reclassified, three more overrides mid-writing. No second coder, no blind re-check, and they admit boundary papers exist in every category. So when they call foundations smallest by far at sixty papers, how much of that survives somebody else doing the sorting? 27 00:13:14,472 --> 00:13:48,930 [Dr. Ada Shannon] Foundations probably survives fine — sixty isn't doubling under recoding. What worries me more is self-evaluation versus training-time self-iteration, since a process-reward paper could land in either bucket depending on which sentence the keyword rule keys off. And there's inflation stacked on that: about fourteen percent of the supplemental harvest, roughly fifty-four papers, is flagged as peripheral bleed by the authors themselves, and it stays in the counts anyway. So that eighty-two-percent-posted-in-2026 figure they use to call self-evaluation fastest-consolidating — 28 00:13:48,930 --> 00:13:55,803 [Hal Turing] Oh wait, hold on — if you strip the bleed back out, you might not have much of a growth curve left, just noise wearing a trend line. 29 00:13:55,803 --> 00:14:24,317 [Dr. Ada Shannon] Except we don't get to find out, because the actual growth stats in Figure Six come from the seed corpus only, not the contaminated supplement — that number's probably clean. But it shows how carefully you have to read this paper, source versus decoration. And it compounds structurally: the corpus is arXiv-only, and they admit a publication-censoring effect hiding exactly the frontier-lab practice everyone worries about. A paper-count survey can describe where public attention sits. It can't tell you where the field's real mass is. 30 00:14:24,317 --> 00:15:06,438 [Hal Turing] That censoring gap is what makes me suspicious of how much weight the Anthropic essay carries here. The paper says it's using Anthropic's stage vocabulary as a motivating frame, not evidence — fine. But A-Evolve-Training gets measured against Anthropic's own scenarios, and 'closing the loop' becomes the taxonomy's terminus. Even the harness-evolution material leans that way — Zhang, Hu, Lu, Lange, and Clune's 2025 Darwin Gödel Machine gets used mainly as a taxonomy anchor, not examined on its own validation loop. Doesn't a lab with a stake in the RSI story end up shaping which findings get the spotlight? 31 00:15:06,438 --> 00:15:42,801 [Dr. Ada Shannon] Partially, yes, and it shows up somewhere unexpected — citations. They say citation counts must not be read as impact, since most of the corpus is 2026 with near-zero citations. Then they call ReST-MCTS* 'the thread's most-cited work' and ScienceAgentBench 'the most-cited paper' to justify foregrounding them. That's using a metric they've already disclaimed. It also raises a question the paper never answers: was any LLM used to classify or triage twelve fifty papers? If so, that classifier inherits exactly the blind spot section five warns about. 32 00:15:42,801 --> 00:16:06,857 [Hal Turing] Which loops back to governance — Anthropic's three scenarios only matter if you can verify which one you're in, and the field has basically no infrastructure for that. The proposals on the table are compute thresholds on training runs, but sections three-five and six-one both say the real gains often come from scaffolding and search engineering, not the base model. A regime watching training compute could miss the actual lever. 33 00:16:06,857 --> 00:16:39,550 [Dr. Ada Shannon] That scaffolding point is the blind spot no draft regulation accounts for yet. It connects to the other fight still live — collapse. Shumailov and colleagues' Nature result stays the pessimistic anchor; the optimistic camp says it's engineerable with grounding and gating. This survey leans on the pessimistic pole while admitting the field keeps producing counterexamples. Their answer is evaluator co-evolution — the judge improving alongside the policy — but that's a bet, not a result. It's the toolbox-versus-takeoff question again, and nobody's shown which. 34 00:16:39,550 --> 00:17:02,956 [Hal Turing] And that mismatch is stark — sixty of twelve fifty studying the stakes everyone claims to care about, against three ninety-three doing deployment engineering. They chalk it up to funding and legibility. Fair enough, but nobody in this corpus measures what happens to the humans doing the auditing once they're pushed from labeler to, in their own words, auditor of the auditor. 35 00:17:02,956 --> 00:17:29,241 [Dr. Ada Shannon] Which is a fair place to land. Strip away the taxonomy and the real thesis is that self-improvement is only as real as its verification. Everything that worked here — FunSearch, the code-repair loops, the training pipelines — worked because something outside the model could say yes or no. Everything that broke, broke where that signal thinned out. The open problems they leave behind are really one problem wearing five hats. 36 00:17:29,241 --> 00:17:47,678 [Hal Turing] That's a solid place to leave it. The interesting question was never whether AI can improve itself — plenty of narrow loops already do, quietly, in production. It's whether anything can verify the improvements that matter most. Thanks for listening, everyone — this has been Hal Turing and Dr. Ada Shannon. Take care.