1 00:00:01,000 --> 00:00:46,450 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into a paper called Agentic Hardware Design as Repository-Level Code Evolution. It's from Cunxi Yu et al. — four co-authors in total, all out of NVIDIA Research — and it hit arXiv on June 26th, 2026. The core question they're asking is a big one: can hardware RTL design be managed the same way software engineers manage a codebase — a design that evolves commit by commit inside git, with an LLM agent doing the editing and a hands-free evaluator deciding what actually gets kept? 2 00:00:46,450 --> 00:01:19,900 [Dr. Ada Shannon] What pulled me in wasn't the benchmark numbers — we'll get to those — it was the framing underneath. Most 'AI writes hardware' papers really just mean 'we fine-tuned a code model on more Verilog.' This one makes an actual structural claim: that repository-level self-evolution, the thing already reshaping how software gets written, generalizes cleanly to an artifact as unforgiving as a chip design. That's a testable idea, not just a bigger benchmark score. Given how conservative hardware teams are about letting anything autonomous near a real design, that's the part worth taking seriously first, before we even look at pass rates. 3 00:01:19,900 --> 00:02:05,775 [Hal Turing] So let's ground why RTL is such a brutal test case in the first place. When an LLM writes you a Python function that's ninety-five percent right, that's usually fixable — patch it, ship a hotfix. A hardware module that's ninety-five percent right isn't ninety-five percent correct hardware, it's just wrong. It might compile, might even simulate fine for a while, then fail on a bit-width edge case or a reset timing quirk or a ready-valid handshake nobody wrote down. And once that design tapes out into silicon, there's no hotfix — you're looking at a multi-million-dollar respin. RTL also runs concurrently on clock edges, not top to bottom like software, so a lot of what a spec assumes is implicit convention the model has to infer. 4 00:02:05,775 --> 00:02:46,824 [Dr. Ada Shannon] Right, and that's exactly the gap HORIZON is built to close. Instead of one-shot generation — spec in, Verilog out, hope for the best — they wrap the whole thing in a loop. You start with a Markdown 'harness': a document describing the objective, relevant domain knowledge, an executable evaluator that can actually compile and simulate the design, and an acceptance predicate — the bar a candidate has to clear. A bootstrap agent compiles that harness into a 'project pack,' basically a scaffolded git repository. A second, hands-free agent then works inside an isolated git worktree — edits files, runs the evaluator, commits if it passes, reverts and retries if it doesn't. Git isn't a changelog here, it's the state manager and the trace of the whole search. 5 00:02:46,824 --> 00:03:37,875 [Hal Turing] And that's not a totally new trick — it's got a lineage. It traces back to AlphaEvolve, the coding agent paper from Novikov and colleagues at Google DeepMind in 2025, which showed pairing an LLM with a real automated evaluator in an evolutionary loop could improve algorithmic kernels — matrix multiplication routines, that kind of thing — well beyond single-shot prompting. Cunxi Yu, the first author here, scaled that up with SATLUTION in 2025, evolving entire SAT-solver repositories under the same propose-evaluate-keep loop instead of isolated functions. Then ABCEvo in 2026, same author, applied repository-scale self-evolution to ABC, the logic-synthesis tool chip designers actually run day to day. HORIZON is— 6 00:03:37,875 --> 00:03:53,699 [Dr. Ada Shannon] Wait, wait, hold on — sorry to cut in — so ABCEvo evolves the tool chip designers use to build things, but HORIZON goes one layer deeper and evolves the actual hardware artifact under construction? That's the jump they're making here? 7 00:03:53,699 --> 00:04:47,974 [Hal Turing] Exactly — that's the missing rung on the ladder, and it's why they frame it as extending repository-scale self-evolution from EDA software to hardware-design artifacts themselves. HORIZON isn't the only thing chasing better RTL generation, so it's worth placing it on the map. One camp is domain-adapted generators — models fine-tuned specifically on Verilog, like RTLCoder, NVIDIA's own ChipNeMo, or the more recent ScaleRTL — still fundamentally one-shot, just better weights behind a single pass. The other camp is iterative or agentic repair systems — AutoChip, RTLFixer, VerilogCoder, MAGE, ACE-RTL — which do loop generation against feedback, closer in spirit to HORIZON. Its bet is that the git-native, repository-level framing is the more general, more auditable way to run that loop. 8 00:04:47,974 --> 00:05:26,849 [Dr. Ada Shannon] And that loop needs somewhere rigorous to prove itself, which is where CVDP comes in — the Comprehensive Verilog Design Problems benchmark. It's seven hundred eighty-three problems across thirteen categories, spanning agentic and non-agentic RTL generation and verification tasks. It exists because the older standbys — RTLLM and Verilog-Eval — were starting to saturate; models topping out on them stops telling you anything useful about where the real gaps are. CVDP is built broader and harder: code completion, testbench stimulus, checker and assertion generation, debugging — the messy variety an actual verification engineer deals with, not just 'generate one module from a clean spec.' 9 00:05:26,849 --> 00:05:48,274 [Hal Turing] Okay, CVDP is the obstacle course, got it — but before we get to numbers, I want the plumbing. You keep saying a harness 'becomes' a project pack. What actually happens in that conversion? Right now that phrase could mean anything from a config file to a fully autonomous system, and I don't think our listeners should just take 'becomes' on faith. 10 00:05:48,274 --> 00:06:20,674 [Dr. Ada Shannon] Fair. A bootstrap agent compiles the harness into the project pack, which has five parts: the agent's policy and tool contract, the executable evaluator itself — compile, simulate, check coverage or assertions — the acceptance predicate that decides pass or fail, a git and runtime policy, and domain skills. After that, it's hands-free — each iteration the agent reads the current worktree, plans an edit, calls tools, produces a patch, and the evaluator decides whether it gets committed or logged as a rejection. 11 00:06:20,674 --> 00:06:32,449 [Hal Turing] So how does it not just repeat the same failed approach on iteration eighty when it already tried that on iteration twelve? Where's the memory living across dozens of these cycles? 12 00:06:32,449 --> 00:07:10,924 [Dr. Ada Shannon] This is the part I actually like — the memory isn't a separate module, it's just git. Staged edits get inspected with git diff, every accepted attempt becomes a commit, and the commit message plus attached git notes carry the evaluator's verdict. git log recovers the whole trajectory. Rejected attempts stay in history too, as negative examples of what didn't work. On top of that, they reuse a persistent model session across iterations, so the harness, the project pack, and stable source get served from prompt cache instead of resent every turn. The newly billed tokens each round are basically just the diff, the evaluator output, and the response. 13 00:07:10,924 --> 00:07:19,524 [Hal Turing] Alright, so what does the actual experiment look like — same backbone model the whole way through, or does it swap models per benchmark? 14 00:07:19,524 --> 00:07:40,750 [Dr. Ada Shannon] They lock it down hard. GPT-5.3 is the single fixed agent backbone for every experiment — no fine-tuning, no per-suite prompt tuning. It runs across ChipBench, RTLLM-2.0, Verilog-Eval, and CVDP categories CID 002 through 016 — completion, spec-to-RTL, modification, reuse, linting, stimulus, checker and assertion generation, and debugg— 15 00:07:40,750 --> 00:07:53,250 [Hal Turing] Wait, hold on — one frozen backbone across that entire range of task types, zero task-specific tuning, and it's one hands-free run per suite? Not best-of-N, not cherry-picked? 16 00:07:53,250 --> 00:08:38,774 [Dr. Ada Shannon] One run, hands-free, per suite. And the headline in Table 1 is blunt: 100% pass rate on every single suite. The only miss anywhere is one ChipBench task, and they trace that to a specification-harness defect in the original benchmark, not an agent failure. Now, the iteration-zero aggregate — the state after just one agentic step — is 47.8%, and they're explicit that's not a standalone Pass@1 score, it's the repo after one pass through the loop. The spread underneath that average is the real story: RTLLM and Verilog-Eval converge in one or two iterations. CID 002, code completion, takes 82. CID 004, modification, takes 36. CID 013, checker generation, starts at the worst first-iteration rate of any category — 3.8% — but then climbs almost linearly to 100% in just 19 iterations, no plateau at all. 17 00:08:38,774 --> 00:08:48,074 [Hal Turing] So a brutal starting point doesn't necessarily mean a brutal path — CID 013 proves that. What does all that grinding cost, though? 18 00:08:48,074 --> 00:09:31,774 [Dr. Ada Shannon] That's Table 2, and it's where the paper gets honest about cost as the real bottleneck. Total run is 210 million tokens, and the nine CVDP categories eat 97.1% of that — CID 002 alone is 26.7% of the entire budget. But about 91% of all tokens are cached input, which is that session-reuse design paying off. Then Table 3 adds a wrinkle on the verification side: CID 012 hits 100% pass rate at iteration 32, but average coverage at that point is only 97.9%, not 100. That's not a shortfall — the acceptance gate stops the moment the CVDP harness passes, full stop. Coverage is just something they measured afterward to see how thorough the passing tests happened to be. It was never the thing being optimized. 19 00:09:31,774 --> 00:10:03,750 [Hal Turing] Okay, so the cost story makes sense. But here's what's bugging me now that we've walked through the wins — Table 1 is one run. One HORIZON run per suite, with GPT-5.3, which is a stochastic model, sampling different completions every time you call it. No repeated trials, no seeds, no error bars anywhere. CID 002 needed eighty-two iterations to converge — if you ran that exact same setup again tomorrow, is it eighty-two again, or is it forty, or 150? We genuinely don't know from what's in this paper. 20 00:10:03,750 --> 00:10:40,549 [Dr. Ada Shannon] And that's a real gap, because agentic repair loops are notoriously run-dependent — small differences in which failure the model happens to fix first can cascade into wildly different iteration counts. Iteration-zero on CID 002 is 3.2%, so we're deep in a regime where early sampling noise could snowball for dozens of steps afterward. A hundred percent pass rate might turn out to be robust — eventually clearing every task with enough budget is plausible — but the specific shape of that trajectory, the '82 iterations' figure, is one draw from a distribution the paper never characterizes. 21 00:10:40,549 --> 00:11:09,274 [Hal Turing] Which loops back to the backbone question. Every number in this paper comes from one frozen, proprietary model — GPT-5.3. But their own related-work section spends real space on RTLCoder, ChipNeMo, ScaleRTL, open, RTL-specialized models built for exactly this domain. So does the git-worktree scaffolding do the actual work here, or is HORIZON just really good at hiding how much of that 100% is GPT-5.3 being GPT-5.3? 22 00:11:09,274 --> 00:11:51,600 [Dr. Ada Shannon] Exactly the question the paper leaves open. It's plausible the harness and acceptance-gate machinery transfers to a cheaper, open backbone — the framework is explicitly generator-agnostic by design. But it's equally plausible that eighty-two iterations of coherent multi-step debugging on CID 002 needs frontier-level tool use and reasoning a smaller model can't sustain. There's a second missing comparison too: HORIZON is only benchmarked against its own iteration-zero, which the authors themselves say isn't a real Pass-at-one. No head-to-head against ACE-RTL, Deng and colleagues out of NVIDIA, 2026, or MAGE, Zhao and colleagues, 2025 — the actual competing agentic RTL systems they call complementary. 23 00:11:51,600 --> 00:12:10,300 [Hal Turing] Wait, sorry, hold on — that connects to something that bugged me in the limitations section. The acceptance predicate is literally 'does it pass the same harness the agent's been staring at for eighty-two iterations.' That's not just a missing baseline, that's a live reward-hacking risk, isn't it? 24 00:12:10,300 --> 00:12:50,925 [Dr. Ada Shannon] It is, and credit to them, they say so directly in Section 5. Full access to simulator logs and failure traces across dozens of iterations is exactly the setup where a model learns to satisfy the visible test rather than the underlying spec. Their own fix points at SWE-bench, Jimenez and colleagues, 2024, which withholds fail-to-pass and pass-to-pass tests until after the patch is submitted. They even cite the follow-up work, Aleithan and colleagues' SWE-Bench-plus, 2024, and Wang, Pradel, and Liu, 2026, showing leaky test suites inflate reported performance. CVDP has no equivalent hidden-test protocol yet, so nothing here rules out over-solving. 25 00:12:50,925 --> 00:13:26,625 [Hal Turing] Then there's the turnaround problem, which feels like the deepest limitation. Every benchmark they picked has fast simulator feedback. But Cunxi Yu's own prior paper, SATLUTION, 2025, needed roughly two hours across eight hundred parallel nodes just to evaluate one SAT-competition run. Real PPA optimization and physical design is the reward HORIZON never tests, and by their own admission the simple edit-evaluate-repair loop breaks once feedback takes hours instead of seconds. 26 00:13:26,625 --> 00:14:02,925 [Dr. Ada Shannon] Right, and that's really the honest core of this paper. Section 3 says outright that hosting EDA-software evolution and architecture-level design is 'a framework goal rather than a completed empirical validation.' So the title promises repository-level hardware design broadly, but what's measured is a git-scaffolded repair loop, one frontier backbone, on nine CVDP categories chosen because they're cheap and fast to evaluate. The genuine novelty is the git-native scaffolding itself — the evaluator-in-the-loop idea predates it, straight out of AlphaEvolve. What's new is porting that onto hardware artifacts instead of software ones. 27 00:14:02,925 --> 00:14:38,500 [Hal Turing] For anyone building actual chip-design tooling, I'd read this as: the scaffolding is worth adopting for fast-feedback verification work, completion, checkers, assertions, where a simulator answers in seconds. I wouldn't read it as evidence this is ready for synthesis-and-timing-closure loops. Worth noting, too — Nathaniel Pinckney and Brucek Khailany both also show up on NVIDIA's recent Nemotron 3 model work, so this team clearly moves between core LLM research and applying it back to their own hardware flows. 28 00:14:38,500 --> 00:14:54,875 [Dr. Ada Shannon] Which is probably the right note to close on. Genuine progress on convergence, honestly reported costs, and refreshingly candid about what it hasn't proven — backbone dependence, reward hacking, long-turnaround reward are all still wide open questions. 29 00:14:54,875 --> 00:15:02,025 [Hal Turing] Great breakdown, Ada, as always. That's it for this one — thanks for listening, and we'll catch you next time.