AI Post Transformers β€” Episode Companion

Do We Still Need GPUs? Rethinking AI with Matrix-Enhanced CPUs

Authors: Jack Dongarra, Torsten Hoefler, Satoshi Matsuoka Institutions: Univ. of Tennessee Β· ETH Zurich Β· RIKEN Version: CACM revised, June 19, 2026

The Two Architectural Bets

Closing the GPU gap on CPU silicon requires solving two independent problems at once: feeding the compute units fast enough, and giving them dense matrix hardware to feed. This diagram traces both threads from AlexNet (2012) to today's matrix-enhanced CPU proposal.
Bandwidth lineage (HBM) Compute lineage (matrix engines) Convergence point

Prefill vs. Decode β€” the lens that matters

LLM inference splits into two phases with opposite bottlenecks. Everything in this episode hinges on which one you're optimizing for.

Reading this page

Every chart below is tagged measured or projected, mirroring the paper's own β€” unusually honest β€” distinction between real hardware results and spec-sheet extrapolation. Hover any element for detail.

Tabs: Memory unpacks HBM's physical structure. Roofline shows why prefill and decode land in different performance regimes. Benchmarks compares the two test chips against GPU baselines. Scaling visualizes the cluster-level decode result and its energy cost.

On-Package HBM: Stacking Instead of Spinning Faster

Conventional DRAM gets bandwidth by clocking a narrow bus harder. HBM stacks DRAM dies vertically and wires them with thousands of through-silicon vias (TSVs), trading clock speed for interface width.

Interface width vs. measured bandwidth

Bus width (bits) and achieved bandwidth (GB/s) across four memory technologies referenced in the episode.

Roofline View: Why Phase Determines the Winner

The Williams–Waterman–Patterson roofline model (2009) plots achievable performance against arithmetic intensity. Decode sits low on the memory-bound slope; prefill needs to reach the compute ceiling.

A64FX vs. LX2 vs. GPU Baselines

Kimi-K2 (1T params, MoE) at 256K context, run with INT4 weights, INT4/INT8 KV-cache quant, DeepSeek Sparse Attention, and EAGLE-3 speculative decoding.

Workload stack β€” measured vs. borrowed vs. spec-sheet

Not every ingredient in the "orthogonal" acceleration stack rests on the same footing. Hover a cell for its provenance.

48 Bandwidth-Only Nodes β‰ˆ 1 GPU Superchip (Decode Only)

Decode throughput scales with aggregate HBM bandwidth. The paper's strongest β€” and only fully measured β€” headline result: 48 A64FX nodes match one NVIDIA GB200 NVL4 node on decode throughput.

Energy premium

The one fully measured disadvantage: CPU-based serving draws 1.75×–4Γ— the per-user power of a GPU fleet, attributed to an HBM/precision generation gap rather than the architecture itself.

References