AI Post Transformers · Episode Companion

Multi-Agent AI Rewrites a Million-Line Chip Design Tool

Autonomous Evolution of EDA Tools: Multi-Agent Self-Evolved ABC
Cunxi Yu, Haoxing Ren · NVIDIA Research & University of Maryland · April 16, 2026
arXiv:2604.15082v1 ↗ 1.2M lines · C · 4,000+ files Planning Agent + 3 Coding Agents Sonnet 4.5 CEC-gated correctness
Pre-Evolution → Planning → Coding Agents → Correctness Gate
A planning agent coordinates three directory-scoped coding agents. Click a coding agent box to see what it owns. Every proposed change must pass a formal equivalence check before it's ever scored.
Click Flow Agent, Mapper Agent, or Logic Min Agent above to see what each one owns.
The Correctness Gate, Step by Step
Click through the six-step cycle every proposed edit goes through — a mismatch at step 4 kills the change before any quality-of-results score is even computed.
From Isolated Kernels to a Four-Layer Codebase
This paper's lineage: FunSearch's math discoveries, AlphaEvolve's algorithmic kernels, SATLUTION's full SAT solver, and now a 1.2M-line EDA tool. Node size scales with codebase size.
Scale Comparison
Toggle between codebase size (log scale) and number of simultaneous optimization objectives.
QoR Across Configurations (lower is better ↓)
Normalized quality-of-results score. Bar height grows as QoR improves. Hover a bar for the exact value.
Baseline (no evolution) Single subsystem evolved Pair evolved All three evolved
Improvement Heatmap by Benchmark Suite
Percent improvement per metric. EPFL's arithmetic circuits run hottest.
Where the $2,400 Went
Token spend across the whole run. Priming the agents on the codebase dominates cost.
$60–80
per cycle
2–3h
per cycle
87,749
lines of C generated
Decomposing the Headline Number
The 8.3% the abstract leans on is only the last segment. Most of the total drop from vanilla ABC comes from wiring in already-published human work.
Reward Signal vs. Reported Results
The benchmark suite that scores every evolution cycle is the same suite the final numbers are reported on.
Abstract Claims vs. Table & Conclusion Caveats
Toggle between what the abstract foregrounds and what's actually qualified in the fine print.
●"Discovers optimizations beyond human-designed heuristics"
●"Learns new synthesis strategies" — no scope given
●8.3% QoR gain, headline figure
●"First self-evolving framework" for EDA code
▲~17% of the total gap is integrating already-published FlowTune/Orchestration/Map work, not LLM discovery
▲Reward benchmarks (ISCAS/EPFL/VTR/ITC'99) are the exact suite the final score is reported on — no held-out set
▲Single PDK (ASAP7 7nm), synthesis-stage STA — not silicon, not post-layout
▲No classical baseline (Bayesian opt, bandit search) run over the same knobs
▲CUHK's EvoPlace shipped the same idea for placement the same year
Scale vs. SATLUTION
Same first author, same rollback-on-regression DNA — but a much bigger, multi-objective target.

References