1 00:00:01,000 --> 00:00:43,500 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper is ContractHIL-HLS: Contract-Aligned Multi-Agent Workflow with Hardware-in-the-Loop Feedback for HLS Design — first author Jingbo Zhang, et al, four authors total, out of Beijing University of Technology, posted to arXiv July 28th, 2026. Here's the hook: the title crams two trust problems into one workflow — it doesn't trust the model's word that a design works, and it doesn't trust a loose prompt to hold its constraints steady across a multi-step build. 2 00:00:43,500 --> 00:01:16,575 [Dr. Ada Shannon] Here's what actually matters, Hal: an LLM-generated FPGA design can compile clean, pass simulation, and still be garbage the moment it hits real silicon — bad timing, routing congestion, power draw blowing the budget, none of which shows up until real synthesis and place-and-route. So the question driving this paper is whether a workflow can carry a design from a plain-language request to traceable, board-tested evidence — more reliably than just prompting the model directly, or handing it role-play instructions. 3 00:01:16,575 --> 00:01:30,924 [Hal Turing] That framing already tells me this is bigger than a 'can the model write Verilog' story. Ada, walk us through the basics — what is High-Level Synthesis, and why can't we treat hardware design like compiling software? 4 00:01:30,924 --> 00:02:06,900 [Dr. Ada Shannon] Before HLS, hardware designers wrote RTL by hand — Verilog, VHDL — specifying exactly what happens on every clock cycle, closer to writing assembly than software. HLS is the compiler technology that takes C or C++ and lowers it into that cycle-accurate hardware description automatically — you write a for-loop, and the tool decides how to pipeline it and how many multiply-accumulate units to build, guided by pragmas like unroll factors and memory partitioning. Vitis HLS, which this paper targets, is the dominant commercial tool — think of it like a deep-learning compiler, XLA or Triton, versus raw CUDA. 5 00:02:06,900 --> 00:02:17,599 [Hal Turing] Okay, so why hasn't HLS just won outright, the way high-level languages killed hand-written assembly? Is this used at scale, or still mostly academic? 6 00:02:17,599 --> 00:02:49,649 [Dr. Ada Shannon] Both. Vitis HLS and Intel's HLS Compiler are real, heavily used tools — video codecs, low-latency trading — and CERN's hls4ml compiles trained neural nets straight into FPGA firmware for real-time particle-detector triggers, in production today. But it never got its SGD moment — there's no gradient anywhere. Unrolling a loop or partitioning memory is a discrete combinatorial search handled by heuristics, not backprop, which is exactly the gap this paper's attacking from the outside. 7 00:02:49,649 --> 00:03:06,499 [Hal Turing] Good handoff into what this paper says is missing. They point at Chip-Chat, RTLLM, and HLS-Eval — proof models can generate hardware code, but not proof of a workflow holding intent and tool feedback together. What's the real gap, Ada? 8 00:03:06,499 --> 00:03:44,399 [Dr. Ada Shannon] Chip-Chat is the one to know — 'Chip-Chat: Challenges and Opportunities in Conversational Hardware Design,' Blocklove, Garg, Karri and Pearce, NYU, 2023. A human designer conversationally prompted a GPT-4-class model until they taped out a small chip, proving an LLM can produce working hardware through iteration — but the human was the feedback loop, relaying results back by hand. RTLLM and HLS-Eval are newer benchmarks checking whether code passes a testbench — snapshots, not persistent workflows. And role-prompted setups have a real failure mode: a constraint set in turn two can quietly vanish by turn five. 9 00:03:44,399 --> 00:04:03,499 [Hal Turing] I'll push on that — I've seen 'structured' framings before that are just an elaborate way of dressing up a prompt template. If the contract's a JSON blob with fields for interface and constraints, what stops the same drift? A field can go stale exactly like a role's memory does. 10 00:04:03,499 --> 00:04:25,974 [Dr. Ada Shannon] I actually disagree with you there, Hal — it's not just semantics. In a role-prompt setup, the 'coder' agent has to reconstruct intent from dialogue history. Here, every agent performs a typed transformation on the same object: language to contract, contract to HTML, HTML plus evidence to implementation. Drift in a conversation is silent. A missing field in a structured document is visible — to a human, a script, the next agent. 11 00:04:25,974 --> 00:04:34,599 [Hal Turing] Visible if someone's actually checking it. Visibility isn't enforcement — you can stare at a field and still not act on it. 12 00:04:34,599 --> 00:05:13,424 [Dr. Ada Shannon] That's fair — and it's genuinely open whether the pipeline acts on those fields correctly. I'm not claiming the contract enforces itself, I'm claiming it changes what's inspectable in the first place. Concretely: a structured contract here means named fields for interface, constraints, compatibility, optimization intent, validation, and rollback policy, replacing prompt text that just gets repeated and mutated. And instead of splitting agents by conversational role, they split by transformation — a Contract Agent turns language into that structured object, an HTML Agent renders it stable and dual-readable, and a Hardware-in-the-Loop Agent implements and revises the design using it. 13 00:05:13,424 --> 00:05:16,924 [Hal Turing] Noted — and what makes that third agent different? 14 00:05:16,924 --> 00:05:26,574 [Dr. Ada Shannon] That's the Hardware-in-the-Loop Agent, and the term isn't something this paper invented — it's borrowed from a completely different field. Decades old, from— 15 00:05:26,574 --> 00:05:36,224 [Hal Turing] Oh wait wait wait — hold on, that's the control-systems thing, right? Testing real engines and flight surfaces before trusting a controller? 16 00:05:36,224 --> 00:06:21,924 [Dr. Ada Shannon] Exactly — control systems, automotive, aerospace. Instead of validating a controller purely in simulation, you wire in the real physical system so it's tested against real dynamics before deployment. This paper repurposes that for hardware design: instead of trusting the model's estimate that a design 'should' work, the Hardware-in-the-Loop Agent runs it through real HLS synthesis, real Vivado place-and-route, and real board bring-up, then feeds those measurements back into the next round or a rollback decision. Their deployment case is a post-quantum cryptography accelerator — FPGA hardware for the new quantum-resistant standards, ML-KEM for key exchange and ML-DSA for signatures. 17 00:06:21,924 --> 00:06:45,099 [Hal Turing] Smart stress test — crypto hardware doesn't get partial credit for 'mostly correct.' So that's the shape of it: a contract that survives the conversation, an HTML layer that keeps it inspectable, and a loop that trusts the chip over the model's optimism. Which raises the obvious question — does any of this machinery actually move the needle on whether the code works? 18 00:06:45,099 --> 00:07:26,424 [Dr. Ada Shannon] It does move the needle, but let's be precise about where it happens. The Contract Agent doesn't write any code — it takes the natural-language request and slots it into named fields: role and platform, the public interface, hard and negative constraints like no dynamic allocation, no STL, no hidden I/O, compatibility with the starter code and testbench, optimization intent, validation checks, and a rollback policy. The HTML Agent then renders that into stable, labeled sections — tables, anchors, checklists — dual-readable, so a person can inspect it and a script can parse it just as easily. Only then does the Hardware-in-the-Loop Agent touch implementation, and when it revises, it's reading that HTML plus real CSim or CSynth evidence from the last attempt, not the raw prompt. 19 00:07:26,424 --> 00:07:39,624 [Hal Turing] So by the time code actually gets written, the model's working off a structured object, not a conversation transcript it might drift from. Okay — how'd they test whether any of that changes outcomes? 20 00:07:39,624 --> 00:08:21,274 [Dr. Ada Shannon] Three regimes, same 94 locally-executable HLS-Eval tasks, same model endpoint, same Vitis HLS toolchain, five samples each. Direct is close to the original HLS-Eval prompt. Contract swaps that for the structured fields alone, no HTML, no hardware loop. ContractHIL-HLS adds both. Headline: single-sample testbench pass goes from 64.0 percent under Direct to 70.2 under Contract, then only to 70.4 under the full workflow. Five-sample pass tops out at 76.6 percent. Of that roughly seven-point gain, six points show up the moment the structured contract is added — before HTML or hardware feedback ever enters the loop. 21 00:08:21,274 --> 00:09:14,849 [Hal Turing] Which makes the rest of the machinery look like it's barely earning its keep. Table III backs that up in a way I didn't expect — tokens per candidate drop from 3741 under Direct to 3555 under Contract, fine, structured fields beat repeated prose. But ContractHIL-HLS jumps to 4308 tokens, and single-sample Can Synth actually drops, 97.2 percent under Contract to 94.7 under the full loop. You're paying more and synthesizing less reliably for two-tenths of a point. It's not uniform underneath either — the full workflow helps c2hlsc, chstone, pp4fpga, rosetta, polybench, but machsuite sits at flat zero, seventeen tasks, zero passes, at one sample or five. 22 00:09:14,849 --> 00:09:53,799 [Dr. Ada Shannon] No — hold on, I actually disagree with that framing, Hal. Zero out of seventeen on machsuite isn't the contract failing, it's the contract doing exactly what it's supposed to do: showing you where the ceiling is instead of quietly hiding it inside a plausible-looking candidate. Every field in that contract governs how a solution is expressed, not whether the model knows the underlying algorithm. If the model can't produce the transformation machsuite needs, better scaffolding around a wrong attempt doesn't make it right. And the token bump is mechanical too — the HTML handoff and implementation evidence get carried forward every iteration. That's what you're paying for, not extra reasoning. 23 00:09:53,799 --> 00:10:06,299 [Hal Turing] Sure, but then what's actually being claimed? If a chunk of the benchmark is structurally immune to the intervention and the rest of the gain is a couple of points, I want to know what's being sold here. 24 00:10:06,299 --> 00:10:42,549 [Dr. Ada Shannon] That's fair, and it's exactly why the paper doesn't pitch this as a universal code-generation fix — it calls the contract a state-alignment mechanism. It preserves interface, constraints, and checks when the model is already close enough to land a plausible answer, and it exposes, rather than papers over, the families where the model lacks the algorithm outright. It's also why the paper treats PQC differently — not a second benchmark, a board-tested engineering case. The contract sets ML-KEM-512 for key encapsulation and ML-DSA-44 for the signature, scores everything by energy-delay product, and caps the design at two bitstreams. 25 00:10:42,549 --> 00:10:50,399 [Hal Turing] And this one actually ships to a board — what did the hardware loop do once real EDA feedback started coming back? 26 00:10:50,399 --> 00:11:34,949 [Dr. Ada Shannon] First pass builds a single-bitstream, end-to-end monolithic image and logs its runtime, power, and timing straight into the HTML contract — 207.3 milliseconds average text runtime across six messages, 171.3 millijoule-seconds EDP, 1.570 nanoseconds worst negative slack, all 144 BRAM tiles in use. The second pass, informed by that measured evidence, selects a KEM/DSA-XOR split instead — separate sender-side and receiver-side bitstreams. That routes at 52.4 milliseconds, 23.6 millijoule-seconds EDP, timing-legal on both images at 0.769 and 0.171 nanoseconds slack — while the receiver still independently verifies the decrypted message rather than trusting the sender's path. 27 00:11:34,949 --> 00:11:47,474 [Hal Turing] That's not a marginal win, that's a four-times runtime drop with better timing margin on both images. But why keep the split at all instead of just pushing the monolithic design harder? 28 00:11:47,474 --> 00:12:16,999 [Dr. Ada Shannon] Because the rollback policy won't let a revision stick unless it clears three gates: both images functionally correct, both route successfully, and the conservative EDP score actually improves — conservative meaning it sums both images' power as if resident at once, even though only one bitstream loads at a time. That's the concrete case where hardware evidence changes a system-level decision, not just a line of code. The monolithic baseline stays in the record as a measured comparison point — the board just wouldn't let it keep the crown. 29 00:12:16,999 --> 00:12:49,549 [Dr. Ada Shannon] ...even though only one bitstream is ever actually loaded on the fabric at a time. So it's a pessimistic score on purpose — the workflow won't take credit for power savings it can't prove, and a revision only survives if both images are correct, both route successfully, and that harsh conservative EDP number still improves. Real synthesis output, real routed timing, real power reports, gating whether a design decision sticks. Which is exactly why I want to turn the critical lens on the benchmark side now, Hal — the numbers that are supposed to justify building all that machinery in the first place. 30 00:12:49,549 --> 00:13:37,799 [Hal Turing] Let's do that, because something's bugged me since Table III. Contract alone takes single-sample testbench pass from 64.0 to 70.2 percent — that's the real jump. Then you add the entire HTML stage and the hardware-in-the-loop agent on top, and pass at one moves from 70.2 to 70.4. Two-tenths of a point, on five samples per task. In that same table, single-sample synthesis success actually falls, 97.2 to 94.7 percent, while token cost per candidate climbs from 3555 to 4308. So: on the benchmark evidence alone, is the full three-agent stack earning that added complexity, or is the PQC case study doing all the real work of justifying it? 31 00:13:37,799 --> 00:14:13,774 [Dr. Ada Shannon] On the benchmark alone, no. With five samples per task, a two-tenths shift in pass@1 is well inside what you'd expect from resampling the same 94 tasks with a different seed — that's not a result, that's noise dressed up as one. The Can Synth drop is the more honest number in that table, because it's a real regression, not a rounding artifact: adding the HTML handoff and implementation agent measurably increases how often a candidate fails to synthesize at all on the first try. You're paying more tokens for a stage that, on this evidence, makes single-sample reliability slightly worse. 32 00:14:13,774 --> 00:14:36,699 [Hal Turing] Okay, but here's my pushback — if the HLS-Eval story is basically a wash, doesn't the PQC result settle the argument anyway? Real board, real rollback policy, a four-times runtime drop, positive timing slack on both images. That's the hardware loop clearly changing a system-level decision. Why does the benchmark also need to carry the weight? 33 00:14:36,699 --> 00:15:05,149 [Dr. Ada Shannon] No — hold on, I actually disagree with you there, Hal. That's one design. One accelerator, run and scored by the same authors who built the workflow, no second board, nobody outside the group replicating it. You can't generalize a claim as broad as the abstract's 'system- and board-level closure' from n equals one, however clean that one result looks. The PQC case is genuinely compelling engineering — I'm not knocking the work — but it's a demonstration, not validation. Thin benchmark evidence plus one deployment case doesn't add up to proof of the architecture as stated. 34 00:15:05,149 --> 00:15:24,699 [Hal Turing] Fair, I'll meet you there. It proves the mechanism can work when it matters, not that it's necessary in general — and the benchmark can't do that job for it either. Different evidence, different weight classes. Which loops into something else that bugged me: the paper never says what model is actually behind any of this. 35 00:15:24,699 --> 00:16:06,199 [Dr. Ada Shannon] Right — 'same model endpoint' is the exact phrase, no name, no size, no vendor. That matters because the whole contract-alignment story could be model-dependent — a weaker model might gain enormously from having interface and constraints spelled out as named fields instead of buried in prose, while a stronger model that already tracks long context well might show a fraction of that lift. It also colors the DeepSeek V3 row in Table II and that bar chart — that's not their number, it's lifted straight from the original HLS-Eval paper, Abi-Karam and Hao out of Georgia Tech, 2025, zero-shot under that paper's own protocol. The text admits it's 'not a controlled head-to-head comparison,' but it still sits right next to their bars, and a reader's eye reads 'beats the baseline' regardless. 36 00:16:06,199 --> 00:16:40,974 [Hal Turing] The bigger miss for me is who's absent entirely. SAGE-HLS — Khan, Mashnoor, Akyash and colleagues out of the University of Florida, 2025, a syntax-aware, AST-guided approach to this exact problem — is sitting right there in their own Figure 2 taxonomy of related work. They place it on the map and never run it against Direct or Contract in Table II. That's the head-to-head that would actually tell us something. And Chip-Chat, the role-prompt failure mode they argue against, stays a conceptual critique too — never benchmarked, just asserted. 37 00:16:40,974 --> 00:17:19,549 [Dr. Ada Shannon] There's a gap that jumps out specifically because this is crypto hardware. The only validation gate on the PQC design is decrypted-message correctness — did the plaintext come back right. Nothing in that contract schema checks for side-channel leakage, timing or power variation correlated with the secret key. That's a well-known failure mode for automatically synthesized crypto accelerators, and it's categorically different from a functional testbench pass. Worse, the loop is optimizing EDP and timing slack — exactly the signals a side-channel attack would exploit. You could end up with a design that's faster, greener, functionally perfect, and still leaking the key through its own power trace. 38 00:17:19,549 --> 00:18:06,375 [Hal Turing] So the practical read: treat the contract as a reliability and state-alignment tool, not a code-generation upgrade — it won't rescue a model that structurally can't solve a task family, machsuite proved that at zero percent. Where it clearly earns its keep is the PQC-shaped problem, board deployment with real rollback discipline. Future work, per the authors, means more board-level designs and automated evidence checking — I'd add: name the model, test other backends, before calling any of this general. That's ContractHIL-HLS — HLS-Eval evidence thinner than the headline suggests, next to one genuinely compelling board-tested case that can't carry the whole claim alone. Thanks for listening, everybody — that's all for this one.