AI Post Transformers · Episode Companion

Main Trust Issue in FPGA HLS Design Workflow

Paper: ContractHIL-HLS Authors: Jingbo Zhang, Haoxiang Sun, Wenbo Wang, Wenbo Zhang Institution: Beijing University of Technology Posted: July 28, 2026
arXiv:2607.25283 →

A multi-agent HLS workflow that stops trusting two things at once: the model's claim that a design works, and a loose conversational prompt to hold its constraints steady across a multi-step build. This page visualizes the architecture, the benchmark evidence, the board-tested post-quantum crypto case, and the gaps the hosts pushed back on.

Three agents, split by transformation, not by role

Language → structured contract → stable HTML → hardware-validated implementation. Each arrow is a typed transformation on the same object, not a new conversational turn.

A missing field in the structured contract is inspectable by a human, a script, or the next agent. Conversational drift across role-prompt turns is not — that's the distinction Ada draws against Hal's "it's just a JSON blob" pushback.

Contract Agent

Converts natural language into named fields: interface, hard/negative constraints (no dynamic allocation, no STL, no hidden I/O), compatibility, optimization intent, validation checks, rollback policy.

HTML Agent

Renders the contract into stable, dual-readable sections — tables, anchors, checklists — parseable by both a person and a script.

Hardware-in-the-Loop Agent

Implements and revises using real Vitis HLS CSim/CSynth, Vivado place-and-route, and board bring-up evidence — borrowed from control-systems / aerospace HIL testing.

94 HLS-Eval tasks, three regimes, same model endpoint

Direct (near-original HLS-Eval prompt) · Contract (structured fields only) · ContractHIL-HLS (full workflow) — five samples per task.

76.6%
pass@5, full workflow
+6.2pt
of ~7pt gain lands at Contract alone
+0.2pt
pass@1 gain from HTML + HIL on top

Synthesis reliability moves the wrong way

Single-sample "Can Synth" — does the candidate synthesize at all on the first try.

Adding the HTML handoff + implementation agent costs 0.75pt of pass@1 lift for a 4.5-point drop in first-try synthesizability, and ~750 more tokens per candidate.

Per-family pattern: the contract exposes the ceiling, it doesn't raise it

Illustrative per-family trend across the three regimes. machsuite is the one hard number the episode gives directly: 0 of 17 tasks pass, at one sample or five, under every regime. Other rows approximate the described "helps c2hlsc, chstone, pp4fpga, rosetta, polybench" pattern — exact per-family percentages were not stated on air.

Board-tested case: ML-KEM-512 / ML-DSA-44 accelerator

Not a benchmark row — a single engineering deployment, scored by energy-delay product (EDP), capped at two bitstreams.

Rollback policy: a revision must clear all three gates

The split design only keeps the crown because the board says so — not because the workflow prefers less monolithic designs.

Conservative EDP sums both images' power as if resident at once, even though only one bitstream loads on the fabric at a time — the workflow refuses to take credit for savings it can't prove.

Evidence rigor, mapped

Five dimensions the hosts pressed on. Lower = weaker evidence support for the paper's broad "system- and board-level closure" claim.

n=1
board-tested design, same authors, no external replication
—
model name / size / vendor never disclosed ("same model endpoint")
0.2pt
full-workflow pass@1 lift, on 5 samples × 94 tasks

Named in the taxonomy, absent from the table

Figure 2 places these in the related-work map. Table II never runs them head-to-head against Direct / Contract.

What the validation gate doesn't check

The only PQC gate is decrypted-message correctness. The loop optimizes EDP and timing slack — exactly the signals a side-channel attack exploits.

A design can be faster, greener, functionally perfect on the testbench, and still leak the secret key through its own power trace. Nothing in the contract schema checks for it.

References