← All episodes Hand-Written PTX vs WMMA: A Precision-Dependent GPU Speedup

Hand-Written PTX vs WMMA: A Precision-Dependent GPU Speedup

Aug 25, 2026
This episode digs into a paper testing whether hand-written PTX assembly beats NVIDIA's WMMA API for Tensor Core GEMM kernels on an L4 GPU, finding that the answer flips depending on numeric precision rather than holding as a universal rule. The discussion covers the hardware distinction between Tensor Cores and regular CUDA cores, and contrasts the convenience of WMMA against the finer control PTX offers through instructions like cp.async, ldmatrix, and mma.sync. A key thread traces why this matters in practice: quantized LLM serving at INT8 or INT4 shifts kernels from compute-bound to memory-bound, making the precision-dependent payoff of hand-tuned PTX directly relevant to running open-weight models cheaply. The episode also addresses the methodological choice to test on a single GPU, arguing that isolating precision and instruction-set effects requires holding hardware constant rather than spreading across devices. Listeners get a concrete framework for deciding when the extra engineering effort of writing raw PTX is worth it versus when it's wasted work.
Sources:
1. Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4 — Matt J. Borowski, Blazej Osinski, 2026
http://arxiv.org/abs/2608.10103
2. NVIDIA Tensor Core Programmability, Performance & Precision — Stefano Markidis, Steven W. D. Chien, Erwin Laure, Ivy B. Peng, Jeffrey S. Vetter, 2018
https://scholar.google.com/scholar?q=NVIDIA+Tensor+Core+Programmability%2C+Performance+%26+Precision
3. Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking — Zhe Jia, Marco Maggioni, Jeffrey Smith, Daniele Paolo Scarpazza, 2018
https://scholar.google.com/scholar?q=Dissecting+the+NVIDIA+Volta+GPU+Architecture+via+Microbenchmarking
4. CUTLASS: CUDA Templates for Linear Algebra Subroutines — Andrew Kerr, Duane Merrill, Julien Demouth, John Tran (NVIDIA), with ongoing project contributors, 2018 (initial release, actively maintained since)
https://scholar.google.com/scholar?q=CUTLASS%3A+CUDA+Templates+for+Linear+Algebra+Subroutines
5. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré, 2022
https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness
6. Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations — Philippe Tillet, H. T. Kung, David Cox, 2019
https://scholar.google.com/scholar?q=Triton%3A+An+Intermediate+Language+and+Compiler+for+Tiled+Neural+Network+Computations
7. Ansor: Generating High-Performance Tensor Programs for Deep Learning — Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, Joseph E. Gonzalez, Ion Stoica, 2020 (OSDI)
https://scholar.google.com/scholar?q=Ansor%3A+Generating+High-Performance+Tensor+Programs+for+Deep+Learning
8. Dissecting the Ampere GPU Architecture via Microbenchmarking — Wei Sun, Ang Li, Tong Geng, Sander Stuijk, Henk Corporaal, 2022
https://scholar.google.com/scholar?q=Dissecting+the+Ampere+GPU+Architecture+via+Microbenchmarking
9. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning — Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, Arvind Krishnamurthy, 2018 (OSDI)
https://scholar.google.com/scholar?q=TVM%3A+An+Automated+End-to-End+Optimizing+Compiler+for+Deep+Learning
10. Understanding Latency Hiding on GPUs — V. Volkov, 2016
https://scholar.google.com/scholar?q=Understanding+Latency+Hiding+on+GPUs
11. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration — J. Lin et al., 2024
https://scholar.google.com/scholar?q=AWQ%3A+Activation-aware+Weight+Quantization+for+LLM+Compression+and+Acceleration
12. FlashInfer: Kernel Library for LLM Serving — Z. Ye et al., 2024
https://scholar.google.com/scholar?q=FlashInfer%3A+Kernel+Library+for+LLM+Serving