1 00:00:01,000 --> 00:01:07,780 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're doing something a little different from our usual neural-net fare. We're looking at "Platform Architecture for Tight Coupling of High-Performance Computing with Quantum Processors" — the NVQLink paper. First author Shane A. Caldwell, et al., thirty authors in total, out of NVIDIA Corporation working with nine other institutions — Sandia National Labs, MIT Lincoln Lab, Lawrence Berkeley National Lab, Oak Ridge, Pacific Northwest National Lab, UC Berkeley, A*STAR's Institute of High Performance Computing in Singapore, and more. Posted to arXiv in October 2025. And Ada, the number that grabbed me right away: they're claiming a 3.96 microsecond round-trip latency between a supercomputer and a quantum control system. That's faster than the blink of an eye by about a factor of a hundred thousand. 2 00:01:07,780 --> 00:01:40,242 [Dr. Ada Shannon] And that number matters because quantum error correction has a clock it can't miss. If the classical computer babysitting your qubits can't respond fast enough, you don't just get a slow program — you get a wrong answer, because the qubits decohere while they're waiting. This paper is basically asking: can you build the wiring closet between a GPU cluster and a quantum chip's control electronics tight enough that the classical side becomes part of the quantum machine's real-time nervous system, instead of some slow service it phones out to? That's a genuinely different engineering problem than anything we usually cover here. 3 00:01:40,242 --> 00:02:33,740 [Hal Turing] Right, so let's back up, because this is new territory for the show. The paper opens by describing two completely different ways people imagine hooking a quantum processor up to a supercomputer. In the first, the QPU is basically a black box — a specialized node the supercomputer occasionally hands work to and waits on, the way you'd treat a slow peripheral. Latency there barely matters. But the paper argues there's a second regime, where the HPC resource has to be tightly coupled to the QPU's own control system to run workloads the QPU needs just to function — most notably quantum error correction decoding. In that regime, the classical compute isn't a peripheral anymore, it's functionally part of the QPU. 4 00:02:33,740 --> 00:03:08,477 [Dr. Ada Shannon] And that's the regime this paper actually lives in. Travis Humble, one of the thirty co-authors here, literally co-wrote one of the earliest papers framing a QPU as an HPC accelerator back in 2017 — "High-Performance Computing with Quantum Processing Units," with K. Britt, out of Oak Ridge, published in ACM's Journal on Emerging Technologies in Computing Systems. That paper sketched the architecture in the abstract. NVQLink is the decade-later, concrete engineering answer — actual FPGAs, actual NICs, actual measured latency — 5 00:03:08,477 --> 00:03:28,400 [Hal Turing] Oh wait wait wait — hold on, so Humble basically wrote the prophecy in 2017 and now he's back building the temple? That's a nice callback. Okay, but walk me through why reaction time specifically is the thing that matters, because I'd have guessed throughput — just processing data fast enough — was the whole ballgame. 6 00:03:28,400 --> 00:04:10,986 [Dr. Ada Shannon] Throughput matters too, but it's a different failure mode. Throughput is whether the decoder can keep up with the stream of syndrome measurements coming off the chip — miss that and you get backlog that snowballs until the machine stalls. Reaction time is the delay from the last measurement to the correction actually landing back on the qubits, and that delay directly eats into fidelity, because every idle cycle the qubit spends waiting is another chance for it to decohere. The paper makes an analogy to superscalar CPUs: throughput is like instructions retired per second, reaction time is like the latency of one instruction working its way through the pipeline. You can have great throughput and still lose the program to bad reaction time. 7 00:04:10,986 --> 00:04:56,775 [Hal Turing] Okay, that's a useful frame. Let's nail down some vocabulary before we go further, because this paper throws a lot of acronyms at you. QPU is the actual quantum chip — the qubits themselves. QSC, quantum system controller, is the surrounding electronics, usually FPGA-based, that generates the control pulses and reads the qubits out — think of it as the chip's peripheral nervous system. Inside the QSC, the actual firmware doing pulse generation is called the PPU, pulse processing unit. On the classical side, the paper defines a real-time host, or RTH — the GPU or CPU node doing the latency-bounded compute — connected to the QSC over what they call the real-time interconnect. 8 00:04:56,775 --> 00:05:34,159 [Dr. Ada Shannon] And here's a design choice worth sitting with before we go further: they explicitly ruled out just plugging the QSC in as a PCIe card next to the GPU. It sounds like the obvious move — direct bus connection, minimal hops — but PCIe slots don't scale. You've got a fixed, small number of them per machine, and as soon as you want to aggregate dozens or hundreds of PPUs, you run out of physical slots and driver headaches. So instead they routed the whole thing over commodity Ethernet, the same fabric datacenters already use to scale to thousands of nodes. It's a less obvious choice, and it's the one that makes the rest of this architecture work. 9 00:05:34,159 --> 00:05:58,401 [Hal Turing] Which is a very NVIDIA-flavored answer, honestly — network your way out of a scaling wall instead of fighting the bus. The paper even calls today's QPUs "somewhat wild objects," which, fair — decoherence in microseconds, bespoke controllers everywhere, no standard interface. That's exactly the mess a standardized real-time interconnect is trying to tame. 10 00:05:58,401 --> 00:06:49,764 [Dr. Ada Shannon] That's exactly the mess Ethernet is designed to survive, actually. So here's the interesting part of the network design: once you're off the PCIe card and onto a NIC, you get to make a choice most people wouldn't expect — they deliberately picked an unreliable link. No retransmission, no acknowledgment handshake with the QSC. Because in a reliable connection, a dropped packet means a retry after some timeout, and that retry jitter is far worse for a real-time control loop than just occasionally losing a tiny packet cleanly. And they lean on RDMA plus NVIDIA's DOCA GPUNetIO library to keep the host CPU and host memory completely out of the loop — the NIC hands data straight into GPU memory, the GPU issues commands straight back to the NIC. No detour through a kernel network stack, no context switch, nothing sitting in a queue waiting for the OS scheduler to get around to it. 11 00:06:49,764 --> 00:07:09,593 [Hal Turing] Okay, unreliable-on-purpose is a genuinely counterintuitive design call, I like that. But it only works if the physical link is clean enough that drops are rare to begin with, right? So walk me through what they actually built to test this, because 'we chose Ethernet' is a claim and 'here's a number' is a measurement — those are different things. 12 00:07:09,593 --> 00:07:56,823 [Dr. Ada Shannon] Right, and the proof-of-concept is refreshingly minimal — one AMD RFSoC FPGA standing in for the QSC, one NVIDIA ConnectX-7 NIC, one RTX PRO 6000 Blackwell GPU as the Real-time Host, all talking over 100 gigabit RoCE — RDMA over Converged Ethernet. The test itself is elegant: the FPGA generates a tiny 32-byte packet stamped with a nanosecond-precision PTP timestamp, fires it across the wire, the GPU on the host side does nothing but loop it straight back — no processing, just catch and return — and the FPGA measures the round trip by comparing its own clock against the timestamp it embedded. It's basically a network ping test, but built for a domain where microseconds are the whole ballgame. 13 00:07:56,823 --> 00:08:23,479 [Hal Turing] And the numbers — sample max of 3.96 microseconds, mean 3.839, median 3.84, standard deviation around 35 nanoseconds. That's tight. Like, that jitter number is almost more impressive to me than the mean, because for a control loop, predictability matters as much as speed — you can design around a slow-but-consistent latency, but a wildly variable one wrecks your scheduling. 14 00:08:23,479 --> 00:09:12,381 [Dr. Ada Shannon] Exactly, and I want to be precise about what that number actually is, because it's easy to let it do more work than it's earned — this is a round-trip network loopback test. There's zero decoding computation happening anywhere in that loop. The GPU catches a packet and throws it back, full stop. That's a legitimate and important number for characterizing the network fabric, but it's not a QEC reaction-time number yet — we'll come back to that gap. On the programming side, this is where CUDA-Q gets extended so a programmer can actually address that whole pipeline from one program. Quantum kernels get the __qpu__ attribute, and inside one you can call device_call to invoke a function on a specific CPU, GPU, or PPU, with the compiler handling the data marshaling automatically. 15 00:09:12,381 --> 00:09:24,223 [Hal Turing] Wait, sorry, hold on — so from inside the quantum kernel itself, mid-circuit, I can just reach out and call a function running on a PPU's firmware, and it looks like a normal function call in my C++ code? 16 00:09:24,223 --> 00:10:03,372 [Dr. Ada Shannon] That's the pitch, yeah — device_call is templated so it works whether you're invoking a plain CPU function or launching an actual CUDA kernel with a grid specification, and there's a companion device_ptr type that gives you a typed handle to memory sitting on any of these heterogeneous devices, so you're not manually tracking whose memory is whose. And critically, there's VPPU — a virtual PPU emulator that implements the same instruction set as the physical hardware, so you can develop and test your real-time protocol against emulated control electronics before you ever touch a physical QSC. That's a real productivity move given how scarce and expensive actual quantum hardware time is. 17 00:10:03,372 --> 00:10:37,458 [Hal Turing] Okay, so now put the network number and the compute side together, because the paper does its own math here and it's a big jump. They cite the Fusion Blossom decoder figures from Camps, Rrapaj, Klymko, Austin, and Wright out of Lawrence Berkeley National Laboratory, published at ISC High Performance in 2024 — 200 teraflops per second to decode a 100-qubit surface-code system, scaling up to 1 petaflop per second at 1000 qubits. Then they pivot to an AI-based decoder estimate and the number balloons. 18 00:10:37,458 --> 00:11:30,446 [Dr. Ada Shannon] It does — they take AlphaQubit, the Bausch et al. decoder out of Google DeepMind and Google Quantum AI, published in Nature in 2024, scale its parameter count up 5x to 25 million parameters, assume a 1 megahertz syndrome cycle rate, tack on a 10x headroom factor for dynamic recompilation and calibration updates, and land at roughly 50 petaflops per second just to decode 100 logical qubits. Now here's the part that made me sit up: Section 6.2 walks through the parallel-window-decoding backlog math from Skoric, Browne, Barnes, Gillespie, and Campbell out of Nature Communications 2023, building on the earlier sliding-window analysis from Chamberland, Goncalves, Sivarajah, Peterson, and Grimberg — Chamberland himself is a co-author on this very paper — and in their worked example they plug in a 20 microsecond round-trip latency, even though this same paper just demonstrated under 4. 19 00:11:30,446 --> 00:12:08,202 [Hal Turing] So the number they use in their own scaling equation is five times worse than the number they measured. That's not a small rounding gap — that's the difference between 'this architecture might just work' and 'this architecture has enormous headroom we haven't even started spending.' Which raises the real question: if the network alone eats 4 microseconds and the compute side wants to run at up to 50 petaflops per second, what fraction of a real, decode-inclusive reaction-time budget does that 3.96 microsecond figure actually represent once you stack an actual decoder on top of it? 20 00:12:08,202 --> 00:12:48,744 [Dr. Ada Shannon] Right, so let's actually answer it instead of just gesturing at it. That's the gap I flagged earlier — zero decoder in that loop, no syndrome, no detector graph, no inference pass. So when you ask what fraction of a real reaction-time budget 4 microseconds represents, the honest answer is: almost none of it, once you add compute. Their own estimate is up to 50 petaflops per second for a 100-logical-qubit AI decoder. Even a fraction of a millisecond of actual inference at that scale dwarfs the network hop. The network number is necessary evidence that the wiring won't be the bottleneck — it's nowhere near sufficient evidence for the whole budget. 21 00:12:48,744 --> 00:13:34,348 [Hal Turing] And it's worth flagging what that setup actually is physically: one RFSoC FPGA, one ConnectX-7 NIC, one RTX PRO 6000 GPU, a single point-to-point link. There's no test anywhere in this paper of what happens when you've got dozens or, at fault-tolerant scale, thousands of PPUs all shoving syndrome traffic onto the same Real-time Interconnect at once. Sub-4-microsecond latency with 35 nanoseconds of jitter is very believable for one clean deterministic link. Whether a switch fabric holds that jitter budget under real incast, with many PPUs bursting syndrome data simultaneously, is a completely open question this paper doesn't even attempt to answer. 22 00:13:34,348 --> 00:14:17,305 [Dr. Ada Shannon] And that ties into something they were surprisingly candid about on the network design itself — remember they chose an unreliable, non-retransmitting link on purpose to avoid retransmission jitter. Their own words are that handling drops is 'left in the hands of the user and software, if necessary.' Fine for a single lost video frame. But a QEC decoder maintains a continuous, stateful record of the stabilizer history. A silently dropped syndrome packet doesn't just cost you one measurement, it can desynchronize the decoder's internal record from ground truth for the rest of the run, and the only mitigation on offer is an optional packet-numbering scheme they don't even mandate. 23 00:14:17,305 --> 00:14:32,119 [Hal Turing] Wait, hold on — optional? For a system whose entire premise is that the classical side becomes part of the quantum machine's real-time nervous system, leaving drop detection as an opt-in feature feels like a genuine gap, not a design nuance. 24 00:14:32,119 --> 00:15:17,863 [Dr. Ada Shannon] It is a gap, and it's one they don't close in this paper. Same pattern shows up in their capacity planning. That 50-petaflop figure comes from scaling AlphaQubit up 5x to a 25-million-parameter model. But elsewhere in this same paper, they cite that exact work as evidence AI-based decoders are hard to scale to large code distances. So they're sizing hardware for fault tolerance using a scaling law from an approach they simultaneously doubt scales. Meanwhile the classical alternative, Fusion Blossom, follows a completely different curve, 200 teraflops at 100 qubits to 1 petaflop at 1000. Two decoder families, two different growth laws, and the paper leans on both without reconciling which one actually wins at scale. 25 00:15:17,863 --> 00:16:13,962 [Hal Turing] That AI-decoder-doesn't-scale-but-we're-using-it-to-size-the-hardware tension is the kind of thing that should have gotten flagged in review. And it rhymes with something else that bugged me reading the runtime section — device, device_ptr, the traits, compiled_kernel, the executor, the Driver API — nearly every one of those is a template signature with the method body omitted. 'void send' with nothing inside. That's fine for a design sketch, but the paper pitches this as a path to industry standardization for every QPU and QSC builder, while the actual reference stack is CUDA-Q, DOCA GPUNetIO, ConnectX-7 — all NVIDIA — with the author list overwhelmingly NVIDIA employees. That's closer to NVLink extending into the QSC than an open standard. 26 00:16:13,962 --> 00:17:44,381 [Dr. Ada Shannon] That's the real scope-versus-claims gap here. The one hard number they have validates the lowest layer of the stack — raw packet transit time in a loopback. Everything above that — the programming model, the trait system, multi-PPU behavior, real decoder integration — is proposed, not demonstrated, and they say so themselves, calling this a 'preliminary specification' meant to invite feedback rather than a ratified spec. That's honest framing, but it means a QSC vendor with years of sunk cost in a closed, proprietary PPU toolchain has no measured switching cost to weigh against adopting this — and no outside vendor has committed publicly yet. Where it's genuinely useful is as an open FPGA IP core that lets a builder skip an HTTP control plane without buying into the whole trait system. Worth noting Ang Li, one of the Pacific Northwest National Lab co-authors here, has also been working on trillion-parameter agentic systems and context compression — a reminder how much this HPC-for-quantum problem overlaps with mainstream ML infrastructure people. The real future work is multi-PPU validation, closing the gap between this network number and a true decode-inclusive reaction time, and calibration workloads as the second real-time use case beyond decoding — all still ahead of us, with today's applied QEC still nowhere near the roughly hundred-logical-qubit regime this architecture is built for. 27 00:17:44,381 --> 00:18:17,260 [Hal Turing] So the takeaway: NVQLink is a genuinely clever, well-engineered answer to one real question — can commodity Ethernet hit microsecond-scale determinism — and the answer there looks like yes. But the paper's ambitions run well past its evidence, and the parts that matter most for whether this actually works at fault-tolerant scale, multi-PPU contention, dropped-packet correctness, and which decoder family to even build for, are still open. That's it for this one — thanks for listening, and we'll catch you next time.