The Agent / Harness / Evaluator Loop
Every self-improvement claim in the 1,250-paper corpus decomposes into three parts: an agent that acts, a harness that can be edited (even by the agent itself), and an evaluator that maps outputs to a quality signal. What differs is whether that evaluator stays fixed.
The Verification Hierarchy
Improvement strength tracks how trustworthy the judge is — formal verifiers at the top, intrinsic self-assessment at the bottom. Hover a tier for examples.
Mapping 1,250 Papers
Two-stage harvest: 871 seed papers across seven threads, then 379 supplemental papers filling taxonomy gaps. Foundations — theory and safety — is the smallest bucket by far.
The Mirror Loop Experiment
Ten rounds of pure ungrounded self-critique across three model providers and four task families. Informational change declined 55% — the models were reformulating, not improving — until one verification step was inserted.
From Good's 1965 Essay to the 2026 Survey
The field traded Schmidhuber's unattainable mathematical proof for empirically checkable but weaker signals. Click a node for details.