AI Post Transformers · Visual Companion

NVQLink: Wiring Supercomputers Directly to Quantum Chips

Platform Architecture for Tight Coupling of High-Performance Computing with Quantum Processors — Caldwell, Khazraee, Agostini, … Humble (30 authors, NVIDIA + nine institutions), October 2025
arXiv 2510.25213 3.96 µs max round trip 100 GbE RoCE CUDA-Q · DOCA GPUNetIO

Peripheral vs. part of the machine

Toggle the two ways to attach a QPU to a supercomputer.

From abstract sketch to measured system

Hover the dots. The 2017 accelerator framing meets a 2025 measurement.

Who built it

NVIDIA plus nine institutions, thirty authors.

Reaction time: last measurement → correction lands

Every microsecond of waiting is idle time for the qubit. Illustrative exponential-decoherence model.

Model is schematic: retained = exp(−t / T_coh). Coherence time and decode time are slider inputs, not paper values.

Throughput is a different failure mode

If decoding each round takes longer than the syndrome cycle, backlog snowballs until the machine stalls.

Cycle time fixed at 1 µs (1 MHz syndrome rate). Backlog = n·(1 − 1/r) rounds when r > 1.

The building blocks

Hover or tap any block. Dashed flows are the real-time interconnect.

Hover a block to see what it does.

Why Ethernet instead of a PCIe card

Slide the number of PPUs. A host has a handful of slots; a switch fabric keeps scaling.

Slot count (4) is illustrative, not from the paper.

One program, many devices: device_call

From inside a __qpu__ kernel, call a function on a CPU, GPU or PPU. Swap the PPU for its emulator.

Round-trip distribution

Mock sample (n = 2000) drawn to match the reported mean, spread and max.

Against the yardsticks

The paper's own scaling example assumes 20 µs.

Anatomy of the loopback test

Step through one 32-byte packet.

Unreliable on purpose

Retransmission jitter versus a cleanly lost packet. Schematic, not measured.

How much classical compute per decoder?

Two decoder families, two growth laws. Toggle series and the AI-decoder scaling factors.

x-axes differ by definition: Fusion Blossom points are qubits of a surface-code system, the AI point is 100 logical qubits. The 1 PFLOP/s baseline is implied by 50 ÷ (5 × 10).

Evidence map: what is shown vs. proposed

Editorial assessment from the discussion. Cold blue = no evidence, red = demonstrated. Hover a row.

Hover a row for the caveat.

References