AI Post Transformers / Visual Companion

When AI Builds Itself and Recursive Self-Improvement

The central split is execution versus direction. The visuals below show why long-horizon engineering loops are clearly tightening while research taste, trustworthy evaluation, and governance still lag.

Source article: Marina Favaro et al., Anthropic Institute, 2026. This is an institute essay rather than a peer-reviewed paper, so the page keeps public benchmark evidence and internal lab evidence visibly separated.

Loop Compression Signal
2026 Signal Balance
>80%
Merged production code attributed to Claude, per Anthropic's May 2026 claim.
76%
Open-ended coding-session success on Anthropic's hardest internal engineering bucket.
3x -> 52x
Fixed-goal research engineering speedup on the same task from May 2025 to April 2026.
51% -> 64%
Model beats the human's next-step choice more often in Anthropic's detour-moment study.

The ladder from tools to self-improvement

Step through the stages. Execution rises earlier than research judgment, which is why the episode treats today's evidence as a tightening loop, not proof of full recursive self-improvement.

Selected stage
The lower lines track three dimensions on a normalized 0-100 scale: execution endurance, research judgment, and self-feedback leverage.

How the loop tightens

Recursive self-improvement needs more than code output. It needs an AI system that can help choose, run, review, and feed those results into the next model cycle.

Dashed segments mark the gap between current strong execution and the harder jump to trusted successor-building.

Execution is hot; problem selection is not

Each row is a source family. Each column is a capability. Hover any cell to inspect what the source is actually saying rather than reading a raw number as a universal score.

Values are source-grounded mock intensities for visualization. They align with the episode's framing: defined loops are getting easier to compress than open-ended science or verifiable slowdown.

Where current systems sit

The key frontier question is whether points migrate up and right together. Today's strongest public and internal signals move rightward first: longer execution, more tool use, more leverage.

Bubble size represents how much a system can feed its work back into future model development.

Public benchmarks versus internal lab claims

Toggle the lens. Both modes use the same visual scale for readability, but the underlying units differ, so compare shapes and gaps, not absolute bar heights.

Current lens
The public lens leans on METR, RE-Bench, CORE-Bench, PaperBench, SWE-bench variants, and the open-source productivity RCT. The internal lens follows the Anthropic essay's 2025-2026 numbers.

Trend shape matters more than any single metric

Public evidence shows growing endurance and partial tool reliability. Internal evidence claims much steeper gains on fixed-goal engineering loops than on research judgment.

The steepest slope belongs to execution. Judgment moves, but slower.

Amdahl's law does not disappear

As code writing gets cheaper, the bottleneck shifts. Review, validation, provenance, deployment, and governance absorb a larger share of the total loop.

Bottleneck mix
The point is not that gains are fake. The point is that faster implementation alone does not guarantee better science or safer autonomy.

Why safety and coordination get pulled forward

If execution loops compress faster than verification capacity, the readiness gap widens. That is the governance pressure the episode keeps returning to.

The hottest risks here are judge bias, automation bias, provenance loss, and the difficulty of verifying a real multi-lab slowdown.

References

Compact source deck for the visuals. The episode also cites prior AI Post Transformers episodes, but the core evidence here is the Anthropic essay plus the public benchmark papers below.

When AI Builds Itself
Anthropic Institute, Marina Favaro et al., 2026
anthropic.com/institute/recursive-self-improvement
Measuring AI Ability to Complete Long Software Tasks
Thomas Kwa, Ben West, Joel Becker, Amy Deng, et al. / 2025
arXiv:2503.14499
RE-Bench
Hjalmar Wijk, Tao Lin, Joel Becker, Sami Jawhar, et al. / 2024
arXiv:2411.15114
CORE-Bench
Zachary S. Siegel, Sayash Kapoor, Nitya Nadgir, Benedikt Stroebl, Arvind Narayanan / 2024
arXiv:2409.11363
PaperBench
Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, et al. / 2025
arXiv:2504.01848
Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
Joel Becker, Nate Rush, Elizabeth Barnes, David Rein / 2025
arXiv:2507.09089
SWE-bench Goes Live!
Linghao Zhang et al. / 2025
arXiv:2505.23419
Multi-SWE-bench
Daoguang Zan et al. / 2025
arXiv:2504.02605
SWE-bench Multimodal
John Yang et al. / 2024
arXiv:2410.03859
Collapse of Self-trained Language Models
David Herel, Tomas Mikolov / 2024
arXiv:2404.02305