1 00:00:01,000 --> 00:00:45,774 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. So today we've got a genuinely provocative one on the table: 'Do We Still Need GPUs? Rethinking AI and Scientific Computing on Matrix-Enhanced CPUs.' That's Jack Dongarra et al., three co-authors total — Dongarra, Torsten Hoefler, and Satoshi Matsuoka — out of the University of Tennessee, ETH Zurich, and RIKEN's Center for Computational Science in Japan. This is the CACM revised version, dated June 19th, 2026. Ada, when I saw Dongarra's name attached to this, I sat up — the man basically wrote the book on dense linear algebra performance. 2 00:00:45,774 --> 00:01:13,500 [Dr. Ada Shannon] He did, and normally when someone asks 'do we still need GPUs' it's a hot take with zero receipts. What actually pulled me into this one is that they don't just make the architectural argument and wave their hands — they back at least half of it with a real measured hardware study on two actual chips, not simulations, not a vendor slide deck. And critically, they're upfront about which half is measured and which half is still a projection. That kind of self-imposed rigor is rare in this genre, and it's exactly why it's worth our full attention today. 3 00:01:13,500 --> 00:01:29,500 [Hal Turing] Good setup — and that measured-versus-projected split is going to matter a lot later. But first, for anyone who hasn't thought hard about it: why did GPUs even become the default for AI in the first place? What exactly could CPUs not do? 4 00:01:29,500 --> 00:02:18,650 [Dr. Ada Shannon] Two things, and the paper leans on both. First, raw parallel throughput — GPUs were built to push the same instruction across thousands of pixels at once, which is structurally identical to pushing the same multiply-add across thousands of matrix entries. When Krizhevsky, Sutskever, and Hinton trained AlexNet at the University of Toronto in 2012, part of the breakthrough was that two consumer GPUs could brute-force what would've taken a CPU cluster weeks. Second, memory bandwidth — NVIDIA started shipping high-bandwidth memory on GPUs with the P100 in 2016, because their arithmetic units had gotten fast enough that ordinary memory couldn't keep up. Tensor cores landed the following year in Volta, purpose-built for the low-precision matmuls deep learning actually runs. Throughput, bandwidth, and the right datatypes, all on one chip. CPUs had none of that — they were optimized for latency and branchy general-purpose code. 5 00:02:18,650 --> 00:02:44,075 [Hal Turing] Okay, so that's the historical gap. The paper's whole bet is that gap can close on the CPU side now — through two ingredients: a matrix-multiply engine built into the CPU pipeline, and on-package high-bandwidth memory, HBM. Let's start with HBM. What is that, physically? Because I think people hear 'more bandwidth' and picture some abstract dial, when it's actually a totally different way of building the memory chip. 6 00:02:44,075 --> 00:03:24,775 [Dr. Ada Shannon] It is a different physical thing entirely. Regular DRAM — your DDR5 sticks, or GDDR6 on a graphics card — sits on its own module, wired to the processor through a fairly narrow channel, maybe 64 bits, and you get bandwidth by cranking the clock speed way up. HBM instead stacks several DRAM dies vertically, wires them together with thousands of microscopic through-silicon vias, and sits that stack right next to the processor die on a shared silicon interposer. That gets you a 1024-bit-wide interface instead of 64. SK Hynix demonstrated the first working version at ISSCC back in 2014 — Lee et al — and that's essentially what JEDEC standardized as HBM shortly after. 7 00:03:24,775 --> 00:03:30,100 [Hal Turing] Wait, hold on — a thousand-bit-wide interface, on one memory stack? 8 00:03:30,100 --> 00:04:18,525 [Dr. Ada Shannon] Versus 64 to maybe 256 bits for conventional memory, yeah. That's the entire trick, and it matters most for one specific pattern in AI inference. There are really two phases to running a large language model: prefill, where the model chews through your entire input prompt at once — that's compute-bound, matrix-heavy, the GPU's classic home turf — and decode, where it generates output one token at a time, which means re-streaming the model's weights and its growing cache from memory for every single token. Decode barely does any arithmetic per byte moved; it's starved by whatever memory system is slower. That's the bandwidth-bound versus compute-bound split — some workloads are capped by how fast you can compute, others by how fast you can feed the compute units. HBM is squarely an answer to the second problem. 9 00:04:18,525 --> 00:04:54,300 [Hal Turing] And the matrix engine is the other half — the compute-bound answer. That's things like ARM's Scalable Matrix Extension, SME, or Intel's AMX, bolted onto the CPU core itself, plus what the paper calls mixed-precision arithmetic — running different parts of a computation at different numeric precisions, full FP64 where you need the accuracy, all the way down to INT4 or FP4 where you can get away with it. So put a matrix engine and HBM on the same CPU die and, on paper, you've closed both historical gaps at once. 10 00:04:54,300 --> 00:05:31,975 [Dr. Ada Shannon] That's the architectural bet exactly. And it's not purely hypothetical — Fugaku's A64FX, an ARM chip with wide vectors and on-package HBM but no matrix engine at all, topped the Top500 list from 2020 to 2022, and simultaneously led HPCG and Graph500, the benchmarks actually built from memory-bound, irregular kernels rather than clean dense matmuls. So the bandwidth half of this story already has a real, deployed proof point. What it doesn't have yet is a matrix engine in the same package — which is exactly the gap the paper's companion study is designed to close, and where things get a lot more interesting. 11 00:05:31,975 --> 00:06:18,150 [Hal Turing] So that's the historical proof point that the CPU side isn't hand-waving. But HPCG and Graph500 are simulation benchmarks — the paper's real test is AI, and they picked a genuinely brutal workload for it: Kimi-K2, a trillion-parameter Mixture-of-Experts model, run at 256K-token context. That's straight out of the Kimi K2 Technical Report from Moonshot AI, 2025 — this whole empirical case rests on that model. And they don't run it naked, either. It's the full modern acceleration stack: INT4 quantization-aware weights, INT4 and INT8 KV-cache quantization, DeepSeek Sparse Attention, and EAGLE-3 speculative decoding all switched on from the start. 12 00:06:18,150 --> 00:06:54,075 [Dr. Ada Shannon] Which is the right call, methodologically — you want to test the architecture against a competent software stack, not a strawman. But one piece of that stack deserves a flag. DeepSeek Sparse Attention didn't come from K2 at all — it's from DeepSeek-V3.2, DeepSeek AI, 2025. The authors are borrowing it and projecting how it would behave on Kimi-K2. That's a reasonable sensitivity analysis, but it's not the same as measuring DSA actually running on K2. Anyone who wants to judge whether that transfer holds should go read the DeepSeek-V3.2 paper directly and form their own opinion, because this paper doesn't validate it on its own target model. 13 00:06:54,075 --> 00:07:37,425 [Hal Turing] Right, and to test that workload they pick two real Arm chips that isolate exactly one variable each. First is the A64FX — the same Fugaku chip from a minute ago — wide vectors, about 1 terabyte per second of HBM, and critically, no matrix engine at all. That's the control. Second is the LX2, a newly shipping Armv9 chip that adds SME on every core, 4 terabytes per second of HBM, and a rated 240 teraFLOPs of BF16 plus 960 TOPS of INT8 per socket. So one chip proves the bandwidth argument, the other is supposed to prove the matrix-engine argument. 14 00:07:37,425 --> 00:08:11,025 [Dr. Ada Shannon] And on decode, the bandwidth argument doesn't just hold — it's already won. The study measures K2 decode as roughly 80% bound by HBM bandwidth and only about 1% bound by floating-point throughput. The matrix engine is basically asleep during decode; it's all about how fast you can stream weights and KV-cache off memory. And because that's true, decode throughput scales with aggregate bandwidth across a cluster — which means roughly 48 A64FX nodes deliver the same K2 decode throughput as a single — 15 00:08:11,025 --> 00:08:24,675 [Hal Turing] Wait, hold on — sorry to cut you off — you said A64FX, the chip with zero matrix hardware, matches a current-generation GPU node on decode? That's not a projection, that's measured? 16 00:08:24,675 --> 00:09:16,850 [Dr. Ada Shannon] Measured, yes — 48 A64FX nodes against one GB200 NVL4 node, comparable aggregate HBM bandwidth, comparable decode throughput. That's the strongest result in the whole paper, and it's the one built on hardware that's been running since 2020. Prefill is the opposite story, though, and it's where the matrix engine actually has to show up. On prefill the A64FX trails a top GPU by about 47x at peak — no surprise, given it has nothing to do the dense matrix work with. Closing that depends entirely on which GPU you're comparing against: something like 80 teraFLOPs per node if the baseline is a sensibly configured GPU running with sparse attention enabled, or north of 750 teraFLOPs per node if you compare against a maximal, dense, fastest-GPU configuration with no such savings. 17 00:09:16,850 --> 00:09:51,375 [Hal Turing] And the LX2's spec sheet — 240 teraFLOPs BF16, 960 TOPS INT8 — clears both ends of that range. Roughly 3x the lower target, and the INT8 number actually beats even the upper one. So on paper, the matrix-engine half of the thesis is met. But I want to be precise about a word you keep choosing carefully — 'on paper.' The A64FX decode result is measured hardware. The LX2 prefill numbers are vendor spec sheets, not a chip anyone has run this workload on yet. 18 00:09:51,375 --> 00:10:49,475 [Dr. Ada Shannon] Exactly the distinction to hold onto, and it's the paper's own framing, not us reading between the lines — one side is confirmed, the other is a spec-sheet promise awaiting a benchmark. There's one more piece we need before the interconnect question makes sense: serving a trillion-parameter model means splitting it across hundreds of nodes, and there are two ways to do that. Tensor parallelism shards each individual layer across nodes and reconciles the pieces with a collective communication step every layer. Pipeline parallelism instead hands whole blocks of consecutive layers to different nodes and just passes activations forward, stage to stage. And that's exactly where Fugaku's Tofu-D torus becomes a constraint: practical all-reduce bandwidth around 6 gigabytes a second, so tensor-parallel prefill turns communication-bound there. Their fix is pipeline parallelism for prefill on A64FX, wide tensor parallelism once you're on a faster fabric like the LX2's. Fine on its own terms — but it leaves a question hanging. 19 00:10:49,475 --> 00:11:39,625 [Hal Turing] Wait, hold on — sorry to jump in, but that's exactly where I want to push. The 48-A64FX-nodes-equals-one-GB200-node decode comparison is built entirely on aggregate HBM bandwidth — add up 48 nodes' worth, it roughly matches one GPU node's. But that assumes those 48 nodes coordinate for free. You just told me their interconnect struggles with all-reduce at prefill. Does that same weakness cost nothing during decode-time coordination — synchronization, KV-cache routing, the failure-domain overhead of keeping 48 separate boxes in lockstep versus one tightly-coupled GPU node? The paper flags Tofu-D as a problem for prefill parallelism. It never revisits it for the decode-scaling claim doing so much rhetorical work. 20 00:11:39,625 --> 00:12:21,000 [Dr. Ada Shannon] That's the gap, and it's not addressed anywhere here. Honestly it's a symptom of something bigger — nearly every headline number in this piece, the 80% bandwidth-bound decode split, the 47x prefill gap, that 6x-to-750x matrix target range, the 1.75-to-4x energy premium, traces back to one companion study that's 'in preparation for arXiv.' It doesn't exist as something we can open and check — no error bars, no methodology, no workload config. And that bandwidth-versus-compute framing running through the whole piece is a direct application of the roofline model, Williams, Waterman, and Patterson out of Berkeley, 2009 — never once named here. Useful lens, uncredited, and now applied through numbers nobody outside the author list can inspect. 21 00:12:21,000 --> 00:13:06,825 [Hal Turing] That uninspectable-companion-study problem hits hardest on the LX2, because the entire matrix-engine half — the actually novel half of the argument — has zero measured results. 240 teraFLOPs BF16, 960 TOPS INT8 is the vendor peak spec sheet. Real matmul kernels routinely land well under peak — 60, 70 percent if you're lucky. Say the LX2 hits 40 to 50 percent of that claimed peak. Their own balanced-baseline target was 80 teraFLOPs, roughly a 6x uplift over A64FX. Halve the LX2 number and you're right back at that floor. Does 'the matrix-engine half is met' survive that, or is it 'maybe, right at the edge, unverified'? 22 00:13:06,825 --> 00:13:46,725 [Dr. Ada Shannon] And it compounds — the workload is one very favorable case. One model, Kimi-K2, one context length, 256K, wrapped in a stack already shrunk as far as possible: INT4 weights, INT4/INT8 KV-cache quant, DeepSeek Sparse Attention, EAGLE-3. They call that stack 'orthogonal to the hardware question,' but it's also the configuration most favorable to a bandwidth-bound CPU story. And DSA's cross-model projection is one unverified layer stacked on the LX2's unverified spec-sheet peak — neither uncertainty is quantified, and stacked together the real error band around the headline conclusion is wider than either number alone suggests. 23 00:13:46,725 --> 00:14:41,100 [Hal Turing] Same pattern in the energy story. They're honest the CPU draws 1.75 to 4x the per-user power of a GPU fleet, and blame a generation gap — older HBM2-class memory, no FP8 yet — rather than the architecture. Fair diagnosis. But 'HBM3e memory and an FP8 engine are on public roadmaps' is itself a projection about hardware that hasn't shipped, and they don't extend GPU roadmaps that same courtesy. Zooming out further — this whole piece frames CPU versus GPU as the entire design space. It never mentions the TPU, Jouppi et al., Google, 2017 — a purpose-built ASIC path that's been viable for a decade. And their own GPU baseline, the GB200 NVL4, is a coherent CPU-GPU superchip from NVIDIA — the binary they're arguing against is already blurred by the machine they're comparing to. 24 00:14:41,100 --> 00:15:17,774 [Dr. Ada Shannon] So practically: decode is what I'd trust today — measured, reproducible, real hardware. Build a bandwidth-heavy CPU tier for the decode-bound part of serving and you're on solid ground. Prefill is still a spec-sheet promise waiting on an LX2 with a stopwatch. And notice what the conclusion leans on to close the sale — not the energy numbers, the one fully measured disadvantage in the piece, but sovereignty and supply chain. Building GPU-free 'for now' is a legitimate motivation on its own. But reaching for it right after conceding a 1.75-to-4x power penalty reads like an admission the efficiency case alone isn't there yet. 25 00:15:17,774 --> 00:15:55,324 [Hal Turing] Fair note to land on — measured where it's measured, hopeful where it's hopeful, and the paper's unusually upfront about which is which. Decode-bound CPU inference: real, today. Prefill-bound matrix-engine parity: plausible, pending the LX2 actually going on a bench. And the sovereignty argument is worth taking seriously on its own terms, separate from whether the FLOPs line up yet. That's 'Do We Still Need GPUs?' — Dongarra, Hoefler, and Matsuoka's case for accelerated CPUs, and our take on where it holds and where it's still asking for trust. Thanks for listening, everyone — we'll catch you next time.