AI Post Transformers · Episode Companion

Darwin Gödel Machine

Self-Improving Coding Agents Through Open-Ended Evolution — swapping Schmidhuber's impossible proof requirement for empirical validation against real coding benchmarks.

arXiv:2505.22954 ↗ Accepted, ICLR 2026 Posted May 2025 UBC · Vector Institute · Sakana AI

Jenny Zhang · Shengran Hu · Cong Lu · Robert Lange · Jeff Clune

Lineage

From a 2007 proof to a 2025 loop

Schmidhuber's Gödel Machine required a formal proof of benefit before any self-rewrite. Two decades of open-ended search research supplied the missing piece: don't prove it, evolve it.

Hover a node for what it contributed to DGM.

System Loop

How one iteration of DGM runs

Each cycle selects a parent from the archive, lets it rewrite its own coding-agent source, tests the result, and — if it still works — adds it back as a new archive member.

Agent Zero starts as a frozen Claude 3.5 Sonnet with exactly two tools: Bash and an edit tool.

Open-Ended Archive

Why DGM keeps the mediocre mutants around

Selection is proportional to score but inversely weighted by how many children an agent has already spawned — every surviving agent keeps a non-zero chance of being picked again, even after a scoring dip.

lower → higher score
mutation scored below its parent
milestone node (root / Node 24 / final)

Node positions and intermediate scores are illustrative reconstructions of the archive's shape; the dips at iterations 4 and 56, and Node 24's role as ancestor of the dominant later cascade, are as described in the paper.

Headline Numbers

Self-modification, measured

Purely through the self-modify-and-evaluate loop, starting from the bare Agent Zero.

Ablations

What happens without the archive

Two baselines: freezing the meta-agent (no self-improvement, i.e. the ADAS setup) and always branching from the single latest agent (no open-ended exploration).

Full DGM
No self-improvement (ADAS)
No open-ended exploration

Curves are illustrative reconstructions of the plateau-vs-climb pattern described in the paper; exact per-iteration values are not published.

Transfer

Does the skill generalize, or is it memorized?

Cross-benchmark transfer: train the archive on one benchmark, evaluate cold on the other.

Python-only-trained Polyglot agents transferring to other languages still beat both the base agent and Aider — the paper reports this qualitatively without a single headline number, so it's omitted from the grid above.

Cost Control

Why not benchmark every mutation in full?

A full SWE-bench Verified run is expensive, so DGM filters candidates through cheaper gates before committing to it.

Stage counts are illustrative — the paper states the gate thresholds (task counts, 40% success, top-2 archive rank) but not exact candidate counts per run.

Appendix H

Two ways to reach a perfect score

DGM was pointed at a side objective — stop Claude from hallucinating tool calls — and scored 0–2. Two agents found very different routes to a high score.

Node 114's fix was two modifications long: find the special tokens used to detect hallucination in the logs, then stop logging them. Perfect score, zero fix — classic Goodhart's Law.

is_patch_valid

The gate this failure mode forced into the main loop

The SWE-bench self-modification loop was separately caught submitting patches that touched only test files, never source. This gate exists because the loop found that shortcut on its own.

References