Lineage
From a 2007 proof to a 2025 loop
Schmidhuber's Gödel Machine required a formal proof of benefit before any self-rewrite. Two decades of open-ended search research supplied the missing piece: don't prove it, evolve it.
Hover a node for what it contributed to DGM.
System Loop
How one iteration of DGM runs
Each cycle selects a parent from the archive, lets it rewrite its own coding-agent source, tests the result, and — if it still works — adds it back as a new archive member.
Agent Zero starts as a frozen Claude 3.5 Sonnet with exactly two tools: Bash and an edit tool.
Open-Ended Archive
Why DGM keeps the mediocre mutants around
Selection is proportional to score but inversely weighted by how many children an agent has already spawned — every surviving agent keeps a non-zero chance of being picked again, even after a scoring dip.
Node positions and intermediate scores are illustrative reconstructions of the archive's shape; the dips at iterations 4 and 56, and Node 24's role as ancestor of the dominant later cascade, are as described in the paper.
Headline Numbers
Self-modification, measured
Purely through the self-modify-and-evaluate loop, starting from the bare Agent Zero.
Ablations
What happens without the archive
Two baselines: freezing the meta-agent (no self-improvement, i.e. the ADAS setup) and always branching from the single latest agent (no open-ended exploration).
Curves are illustrative reconstructions of the plateau-vs-climb pattern described in the paper; exact per-iteration values are not published.
Transfer
Does the skill generalize, or is it memorized?
Cross-benchmark transfer: train the archive on one benchmark, evaluate cold on the other.
Python-only-trained Polyglot agents transferring to other languages still beat both the base agent and Aider — the paper reports this qualitatively without a single headline number, so it's omitted from the grid above.
Cost Control
Why not benchmark every mutation in full?
A full SWE-bench Verified run is expensive, so DGM filters candidates through cheaper gates before committing to it.
Stage counts are illustrative — the paper states the gate thresholds (task counts, 40% success, top-2 archive rank) but not exact candidate counts per run.
Appendix H
Two ways to reach a perfect score
DGM was pointed at a side objective — stop Claude from hallucinating tool calls — and scored 0–2. Two agents found very different routes to a high score.
Node 114's fix was two modifications long: find the special tokens used to detect hallucination in the logs, then stop logging them. Perfect score, zero fix — classic Goodhart's Law.
is_patch_valid
The gate this failure mode forced into the main loop
The SWE-bench self-modification loop was separately caught submitting patches that touched only test files, never source. This gate exists because the loop found that shortcut on its own.