Self-improvement re-cast as tree search: nodes are agents, edges are self-rewrites, and the question each system answers differently is which node to expand next.
Schmidhuber, who wrote the original 2003 proof, is a co-author on the Huxley-Gödel Machine. The acknowledgments thank Jenny Zhang and Shengran Hu (Darwin Gödel Machine authors) for sharing implementation insights — a documented lineage, not just a citation trail.
Every self-rewrite spawns a child node. The field's default policy picks the next parent by its own raw benchmark score — HGM argues that signal is structurally misleading.
Agent A1 fails 9 of 10 tasks but keeps getting re-tried because it's new. Agent A21 only solved a handful itself — but its descendants A211 and A212 are quietly strong. Scoring by immediate performance starves A21's productive lineage. Toggle the view below.
Computing empirical CMP from finished search trees and correlating it against each method's own guidance signal shows HGM's estimator actually tracks what it predicts.
HGM treats each clade as a bandit arm with an uncertain true CMP. Thompson sampling keeps a belief distribution per arm and samples to pick what to expand next — balancing exploration against exploitation automatically.
As the schedule τ tightens, sampling polarizes toward the leading clade instead of spreading evenly.
SICA context-overflowed and stalled after 360 evals — inside 45% of budget.
Polyglot weighted (0.626) vs unweighted (0.873) diverge more than SWE-Verified-60 — likely quantization noise from the int4/int8 self-mod backbone used on that $5k run.
Held on SWE-Lite (unseen tasks + standard split, beating SWE-agent on the same backbone) and again with a GPT-5 backbone swap — essentially tied with the best officially checked SWE-agent submission.
DGM and SICA couple expansion to evaluation: make a child, immediately test it, no way to bail on one that's clearly cooked. HGM splits that into three separate decisions run asynchronously across every available CPU.
Loose and exploratory early in the budget, tightening and polarizing sampling toward the leaders as the budget runs out.