This page treats ACE as an execution-substrate proposal, not a model change. The visuals focus on the real tension in the episode: ACE makes matrix math denser, but full-kernel speed still depends on packing, memory traffic, tails, format conversion, and how tightly datatype semantics are nailed down.
Whitepaper v1 • 2026-04-15AMD + Intel + x86 Ecosystem Advisory GroupStatus: ISA proposal, not product launchHeadline primitive: 16× compute density for INT8 and BF16ACE Whitepaper PDFarXiv Context
Selected References
[1] ACE Whitepaper
The x86 EAG proposal introducing outer-product matrix updates, 8 tile registers, and AMX palette reuse.
x86ecosystem.org PDF
[2] Power ISA MMA, 2021
Commercial precedent for matrix-native accumulation inside a general-purpose CPU.
Google Scholar
[3] Hello SME!, 2024
Compiler-generated matrix kernels for Arm SME, useful for contrast with ACE software enablement.
Google Scholar
[4] SparAMX, 2025
A shipping-AMX comparison point for CPU-side LLM inference work.
Google Scholar
[5] Compute Or Load KV Cache? Why Not Both?
Parallel compute-plus-I/O scheduling for long-context inference.
arXiv:2410.03065
[6] Towards Fully FP8 GEMM LLM Training at Scale
Why low-precision success depends on architecture plus numeric recipe.
arXiv:2505.20524
[7] SOLE
Reminder that non-GEMM kernels can dominate latency and efficiency.
arXiv:2510.17189
[8] SageAttention3 / NVFP4
Recent low-bit training and inference work showing that format choice is a full-stack problem.
arXiv:2505.11594arXiv:2509.25149
All charts and heatmaps use realistic mock values shaped by the episode discussion and cited paper framing. They illustrate likely behavior and tradeoffs, not published ACE silicon benchmarks.