1 00:00:01,000 --> 00:00:42,425 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're reading "Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI." It's a company technical report from the Architect Labs Team at Architect Labs in Palo Alto, with 27 contributors listed in the appendix and no single lead author. The question: can one AI system, given only a spec from two human architects, generate and verify a whole accelerator fast enough to keep pace with the workloads, and beat a comparable edge GPU on efficiency? 2 00:00:42,425 --> 00:01:03,600 [Dr. Ada Shannon] Two weeks from spec to verified RTL, then a Qwen model running on an FPGA in week three. That's a schedule claim as much as a chip claim. The motivation is a timing mismatch: architectures freeze years before silicon ships, while workloads change in months. So hardware pays twice, first in generality added as a hedge, then when a new workload maps badly onto frozen silicon. 3 00:01:03,600 --> 00:01:40,950 [Hal Turing] They back it with two numbers: only 14% of IC and ASIC projects reach first-silicon success and 75% run behind schedule, while vendors advertise up to 10x productivity from AI. Their thesis is that task-level help, chatbots and script generators, hasn't shortened whole-program schedules. So they collapse architecture, RTL, verification, firmware and kernels into one spec-driven loop. But Ada, I don't think 14% indicts RTL tools. First-silicon failures are mostly physical design and integration surprises. 4 00:01:40,950 --> 00:02:00,800 [Dr. Ada Shannon] I disagree, Hal. Their argument isn't that the tools are bad, it's that the handoffs are the schedule. Architecture waits on the performance model, RTL waits on architecture, verification waits on RTL, kernels wait on hardware. Speed up each box tenfold and the waiting between boxes is untouched. It's Amdahl's law applied to an org chart. 5 00:02:00,800 --> 00:02:10,875 [Hal Turing] Sure, but the statistic is about silicon success, not waiting. Different failure modes, and the paper borrows one to motivate a fix for the other. 6 00:02:10,875 --> 00:02:52,525 [Dr. Ada Shannon] That's a fair hit. Treat the statistic as their stated motivation, not evidence. The single-spec co-design idea stands or falls on what they demonstrate. Here's the landscape they set up. ChipNeMo, NVIDIA, 2023, Mingjie Liu first author, adapted LLMs for engineering chat, script generation and bug triage, which is task-level assistance inside the normal flow. Chip-Chat, Jason Blocklove at NYU, 2023, had an LLM co-design a small accumulator-based processor taped out through Tiny Tapeout. And Cheng and colleagues, "Pushing the Limits of Machine Design," from the Chinese Academy of Sciences in 2023, automated a RISC-V CPU design, validated on FPGA and reported running Linux. 7 00:02:52,525 --> 00:03:05,775 [Hal Turing] Okay, I genuinely want to understand the target. It's single-batch, low-power, low-latency inference for physical AI, so robots and drones. Why is that memory-bound? I'd have guessed compute. 8 00:03:05,775 --> 00:03:46,100 [Dr. Ada Shannon] At batch size one, every generated token streams essentially all the model weights from DRAM once, and each weight is used for roughly one multiply. Tokens per second is set by memory bandwidth, not MAC count. The roofline model says an operator's time is the max of compute time and memory time, summed across sequential operators. It's a ceiling, not a measurement. Redwood is a spatial dataflow accelerator: a mesh of identical tiles where you schedule by placing data and work on the mesh, not through a cache hierarchy and a generic instruction stream like a GPU. Redwood Nano is a 2x2 instance on an FPGA, which is basically a very expensive breadboard for chips. 9 00:03:46,100 --> 00:03:53,050 [Hal Turing] Oh wait wait wait— so when they say Redwood runs Qwen, that's on the FPGA, not silicon? 10 00:03:53,050 --> 00:04:30,800 [Dr. Ada Shannon] Right, hold onto that. The FPGA result is measured, on an AMD Versal VPK180 at 250 MHz. The Samsung 8 nm numbers at 1 GHz are projections, scaled from FPGA profiles with gate-equivalent area and CV-squared-f power models. Verification vocabulary: UVM is the standard simulation methodology, and coverage measures how much code and specified behavior the tests exercised. 95% is a strong signal, but it isn't 100% and it isn't proof. Formal verification proves properties for all inputs; the paper mentions a proprietary formal engine and gives no details. 11 00:04:30,800 --> 00:04:40,700 [Hal Turing] Fine. Now walk me through the hardware, because I have CRV, CMXM, CVXM and CMEM in my notes and can't tell them apart. 12 00:04:40,700 --> 00:05:17,800 [Dr. Ada Shannon] It's an N by M mesh of identical tiles. Each tile has a RISC-V control core, the CRV, a matrix engine, CMXM, doing GEMM and GEMV, and a SIMD vector engine, CVXM. Each tile also has 512 KB of local scratchpad, CMEM. A credit-based network-on-chip moves weights and activations, with edge and global DMA engines talking to DRAM over AXI4. A global control MCU sequences everything, with a 48-bit global timer broadcast to every tile. There's a front-end and back-end split and a task manager called CTM, but that's for later. 13 00:05:17,800 --> 00:05:45,100 [Hal Turing] So, headline claims, stated as claims. Under two weeks from spec to verified RTL, 95% coverage on every block, a third week to bring up Qwen3-0.6B, and respin cycles under 48 hours. The abstract also says 1.75x the throughput of a Jetson Orin Nano, 1.9x lower power, and 3.4x performance per watt, where that last figure is tokens per second per watt. 14 00:05:45,100 --> 00:06:03,075 [Dr. Ada Shannon] And one more: Qwen running on Redwood helped design the next Redwood, which the authors frame as early recursive self-improvement. That's their framing. In the next stretch we'll separate what was measured on the FPGA from what was projected, because those abstract numbers sit on the projected side. 15 00:06:03,075 --> 00:06:10,675 [Hal Turing] Before the numbers, Ada, give me the programming model in one pass. How does a host get a tile to do anything? 16 00:06:10,675 --> 00:06:49,525 [Dr. Ada Shannon] Scheduling lives in software, not hardware. That's the bet. Each tile splits into a front end and a back end, so the RISC-V core stays minimal and can clock slower or gate off. It enqueues tasks with IDs into a task manager that handles fencing, loops and tracing, then idles. Messages over the SoC fabric let those managers wait on each other, so the compiler schedules prefetch and double-buffering instead of hardware arbitrating. The host loads dispatch programs onto the MCU and kernels onto the tiles, then writes a dispatch ID. FlashAttention, the Tri Dao paper out of Stanford in 2022, is one dispatch program launching tile kernels over KV blocks and heads. 17 00:06:49,525 --> 00:07:26,375 [Hal Turing] Now the measurement. On the FPGA there are four 128-bit AXI streams giving 16 GB/s from LPDDR4. Qwen3-0.6B runs INT8, averaged over 128 tokens, including prompt transfer and per-token return. Jetson runs at 1020 MHz with 68 GB/s of LPDDR5, measured through NVIDIA's WebUI. The result: 12.1 tokens a second, about 13 at peak, against 28. So the measured Nano is roughly 2.3 times slower. 18 00:07:26,375 --> 00:08:11,425 [Dr. Ada Shannon] Yes, and that's the only measured head-to-head. Then comes the FPGA roofline. Peak GEMV is 128 Gop/s: four tiles, 64 lanes, 250 MHz, times two. Sustained DRAM is 14.04 GB/s. One decoder layer takes 1.241 ms, dominated by the Q/K/V, O, gate/up and down projections. Add embedding and the LM head and a token is 46.02 ms, or 21.73 tok/s, moving 0.627 GB: 44.65 ms of DRAM service against 12.29 of arithmetic. Fully serialized, it's 17.56. So 12.1 is about 56% of the ceiling, and the paper blames launch, synchronization, pipeline fill and drain, and host overhead. 19 00:08:11,425 --> 00:08:17,550 [Hal Turing] That per-operator table is careful work. Now, how do they get from there to silicon? 20 00:08:17,550 --> 00:08:57,025 [Dr. Ada Shannon] Five steps. One, assume a 1 GHz logic clock on Samsung 8 nm, which they call reasonable given FPGA timing. Two, the DRAM pin rate stays fixed, but tile ingress goes from 16 to 64 GB/s by engaging three controllers. The paper says it assumes the same bandwidth as the Jetson's 68, with tile ingress the narrowest stage at about 64. Three, compute engines scale 4x. Four, that gives an architectural ceiling near 95 tok/s. Five, the most conservative projection, 49, comes from profiling the FPGA run and scaling each step for clock, bandwidth, and better scheduling from fewer hardware restrictions. 21 00:08:57,025 --> 00:09:08,550 [Hal Turing] Hold on, that's actually— 49 over 12.1 is about 4.05. That's the bandwidth multiplier and the clock multiplier, both exactly four. 22 00:09:08,550 --> 00:09:30,525 [Dr. Ada Shannon] I actually disagree with you there, Hal. You can't read the method off a ratio. The text describes scaling a profile step by step, and a sum of scaled steps could land near 4x for plenty of reasons. There's no breakdown by factor in the text, so what the ratio means is something neither of us can settle. What I can add is that 49 is 52% of the 95 ceiling, against 56% of roofline measured on the FPGA. 23 00:09:30,525 --> 00:09:35,675 [Hal Turing] Fair enough. Then area and power, since they feed the efficiency number. 24 00:09:35,675 --> 00:10:35,925 [Dr. Ada Shannon] Area uses a gate-equivalent method: 2 million combinational cells, 500,000 flops, plus 15% for test logic, 70% placement utilization, and 20% for clock tree and timing cells. That gives about 2.88 mm². Power is CV²f at 0.75 V and 1 GHz: 0.958 W dynamic, 0.07 W static, 1.335 W chip-side with the rest of the SoC. They say it aligns with FPGA testing and is an upper bound, since clock and power gating are excluded. Jetson averages 2.59 W for CPU and GPU with fusion enabled, and both figures exclude memory controllers. Table III then reads 49 over 28 as 1.75x, 2.59 over 1.335 as 1.94x, and 36.7 versus 10.8 tokens per second per watt as 3.4x, which is just 1.75 times 1.94. 25 00:10:35,925 --> 00:11:12,675 [Hal Turing] Last piece, the system itself. ALP treats the specification as the single source of truth, with no freeze and zero pre-existing accelerator IP. Repository merges peaked at 115 in a day. Verification is fully automated: UVM, testbenches, SVA and formal artifacts from a first-version proprietary formal engine. Every block hit 95% code and functional coverage, there were zero bugs on the first RTL drop going to FPGA, and no bug has yet escaped simulation into hardware. 26 00:11:12,675 --> 00:11:54,675 [Dr. Ada Shannon] Then exploration and software. The SIMD reduction engine was explored over several days across performance, area, timing and coverage, and the paper says candidates differ in control path, datapath and state machine, not just bit-width. Firmware, kernels and performance models were co-developed before RTL. A custom environment multiplexes one FPGA across hundreds of concurrent agents, cutting optimization runs from about 15 hours to 15 to 30 minutes. Finally, Qwen3 was deployed on Redwood as an inference endpoint, and repeated sampling found timing improvements and kernel optimizations for its own operations at zero inference cost. The paper gives no numbers for those gains. 27 00:11:54,675 --> 00:12:20,000 [Hal Turing] Ada, the abstract's headline numbers are all projected. The measured result, 12.1 against 28 tokens per second, only shows up in Table I and the conclusion. So why is the unmeasured number the headline? And here's what I actually want to understand: the projection needs 64 GB/s into a 2x2 mesh at 1 GHz. Was that memory system ever modeled or emulated? 28 00:12:20,000 --> 00:13:13,650 [Dr. Ada Shannon] None of the abstract's efficiency numbers was measured on Redwood. That's the consequence: a projected numerator multiplies an estimated power figure, all against a measured Jetson. On the memory system, the paper does not report multi-controller contention, LPDDR5 bank and refresh behavior, or edge-DMA saturation, and its own footnote leaves the NoC shape abstract. The 1 GHz clock is 'reasonable given FPGA timing', but a 250 MHz fabric result says very little about standard-cell closure on 8 nm. 'Most conservative' is a statement of intent. The 52% versus 56% means it assumes the same efficiency. But launch and sync overheads don't shrink 4x, and Amdahl says their share should grow as the memory term collapses. The paper also credits 'fewer hardware restrictions', which pushes the other way. Nobody can weigh those without an ablation or an error bar, and the paper gives neither. 29 00:13:13,650 --> 00:13:41,750 [Hal Turing] Sorry to cut you off, but the Jetson side bugs me more. Our own arithmetic: 28 tokens a second times 0.627 gigabytes is about 17.6 GB/s, roughly 26% of its 68. It was measured through the Jetson WebUI, with no runtime, quantization or power mode named. If that ran FP16, it moves twice the bytes per token, and in a memory-bound workload the gap is largely precision, not architecture. 30 00:13:41,750 --> 00:14:20,700 [Dr. Ada Shannon] I actually disagree, Hal. You just asserted FP16, and the paper's silence supports that no more than it supports the opposite. Twenty-six percent could be weak software, a different byte count, or a conservative power mode. What we can say is that it's unknown, and that's the paper's failure. The power scope is worse. Jetson's figure is measured, Redwood's is estimated, and both exclude DRAM. At 49 tokens a second, 0.627 gigabytes per token is about 30 GB/s of traffic. At LPDDR5-class energy, my rough estimate is on the order of a watt, comparable to the whole chip figure. Whether the 3.4x survives that is an open question. 31 00:14:20,700 --> 00:14:43,526 [Hal Turing] Fine, I'll retreat from asserting FP16 and demand the measurement instead. Hundreds of agents tuned Redwood's kernels, so where's an INT8 or W4 TensorRT-LLM Jetson, tuned just as hard? On novelty, the paper says public AI-designed chips are toy cores. Given the earlier work we covered, what's genuinely new here? 32 00:14:43,526 --> 00:15:28,251 [Dr. Ada Shannon] A whole accelerator including kernels and firmware, a claimed no-human-below-spec flow, and a transformer running on the FPGA. That's a real step past ChipNeMo's task-level assistance. 'First' and 'production-worthy' overreach: there's a 2x2 array, no physical design, no tapeout, no silicon. Two weeks covers RTL, verification and bring-up, and the 14% first-silicon statistic lives in the parts left out. Verification is the least evidenced claim. The same system writes the RTL, the tests and the assertions, so blind spots can be shared. There are no bug counts, mutation results or proof counts, and no numbers on VerilogEval from NVIDIA in 2023, CVDP in 2025, or FVEval in 2024. A skeptical verification engineer wants escape rates, spec-to-assertion traceability, and an independently written suite. 33 00:15:28,251 --> 00:15:41,201 [Hal Turing] Genuine question on the 48 hours: is the wall clock agent runtime or FPGA bitstream generation? The paper doesn't say. And what about the recursive self-improvement claim? 34 00:15:41,201 --> 00:16:09,926 [Dr. Ada Shannon] It's best-of-N sampling with a verifier, not a system improving its own capability. KernelBench from Stanford in 2025 shows how often AI kernel speedups are artifacts, and AlphaEvolve from Google DeepMind in 2025 is the better example, with real acceptance procedures. What holds up: the FPGA emulation loop, spec-as-source-of-truth co-design, and a working Qwen3-0.6B with an honestly reported 56% of roofline. The 3.4x is a design-target hypothesis, not a result. 35 00:16:09,926 --> 00:16:37,026 [Hal Turing] So what converts the projection into a result? A multi-controller LPDDR5 emulation, placed-and-routed power and timing, sensitivity across the 4x assumptions, more models and contexts, a tuned Jetson, an independent verification audit, and silicon. The takeaway is an impressive demonstration of engineering velocity, with an efficiency claim that is still an extrapolation. Thanks for listening, and goodbye.