This episode explores AMD’s open-source MIOpen library and why deep learning primitives such as convolution, pooling, normalization, and activations are the layer where model performance meets GPU hardware reality. It explains how CNN throughput depends on low-level execution choices, comparing approaches such as im2col-plus-GEMM and Winograd convolution, and shows why libraries like MIOpen need solver-based algorithm selection and auto-tuning to match different workload shapes, precisions, and GPUs. The discussion also covers mixed-precision support, especially bfloat16, along with kernel fusion and composable kernels as ways to reduce memory traffic and launch overhead while keeping vendor-library speed. Listeners would find it interesting because it turns “invisible infrastructure” into a concrete systems story about how open-source GPU software can shape real model training and inference performance.
Sources:
1. MIOpen: An Open Source Library For Deep Learning Primitives — Jehandad Khan, Paul Fultz, Artem Tamazov, Daniel Lowell, Chao Liu, Michael Melesse, Murali Nandhimandalam, Kamil Nasyrov, Ilya Perminov, Tejash Shah, Vasilii Filippov, Jing Zhang, Jing Zhou, Bragadeesh Natarajan, Mayank Daga, 2019
http://arxiv.org/abs/1910.000782. cuDNN: Efficient Primitives for Deep Learning — Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, Evan Shelhamer, 2014
https://scholar.google.com/scholar?q=cuDNN%3A+Efficient+Primitives+for+Deep+Learning3. MIOpen: An Open Source Library For Deep Learning Primitives — Jehandad Khan, Paul Fultz, Artem Tamazov, Daniel Lowell, Chao Liu, Michael Melesse, Murali Nandhimandalam, Kamil Nasyrov, Ilya Perminov, Tejash Shah, Vasilii Filippov, Jing Zhang, Jing Zhou, Bragadeesh Natarajan, Mayank Daga, 2019
https://scholar.google.com/scholar?q=MIOpen%3A+An+Open+Source+Library+For+Deep+Learning+Primitives4. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning — Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Meghan Cowan, Haichen Shen, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, Arvind Krishnamurthy, 2018
https://scholar.google.com/scholar?q=TVM%3A+An+Automated+End-to-End+Optimizing+Compiler+for+Deep+Learning5. Ansor: Generating High-Performance Tensor Programs for Deep Learning — Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, Joseph E. Gonzalez, Ion Stoica, 2020
https://scholar.google.com/scholar?q=Ansor%3A+Generating+High-Performance+Tensor+Programs+for+Deep+Learning6. Automatically tuned linear algebra software — R. Clint Whaley and Jack J. Dongarra, 1998
https://scholar.google.com/scholar?q=Automatically+tuned+linear+algebra+software7. Fast Algorithms for Convolutional Neural Networks — Andrew Lavin and Scott Gray, 2015
https://scholar.google.com/scholar?q=Fast+Algorithms+for+Convolutional+Neural+Networks8. TVM: end-to-end optimization stack for deep learning — Tianqi Chen et al., 2018
https://scholar.google.com/scholar?q=TVM%3A+end-to-end+optimization+stack+for+deep+learning9. Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions — Nicolas Vasilache et al., 2018
https://scholar.google.com/scholar?q=Tensor+Comprehensions%3A+Framework-Agnostic+High-Performance+Machine+Learning+Abstractions10. oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning Compilation — approx. Intel oneDNN graph compiler authors, 2023
https://scholar.google.com/scholar?q=oneDNN+Graph+Compiler%3A+A+Hybrid+Approach+for+High-Performance+Deep+Learning+Compilation11. SparseTIR: Composable Abstractions for Sparse Compilation in Deep Learning — approx. TVM / SparseTIR authors, 2023
https://scholar.google.com/scholar?q=SparseTIR%3A+Composable+Abstractions+for+Sparse+Compilation+in+Deep+Learning12. Autotuning Convolutions Is Easier Than You Think — approx. tensor-compiler autotuning authors, 2023
https://scholar.google.com/scholar?q=Autotuning+Convolutions+Is+Easier+Than+You+Think13. Haotuner: A Hardware Adaptive Operator Auto-Tuner for Dynamic Shape Tensor Compilers — approx. Haotuner authors, 2024
https://scholar.google.com/scholar?q=Haotuner%3A+A+Hardware+Adaptive+Operator+Auto-Tuner+for+Dynamic+Shape+Tensor+Compilers14. The Case for Training Large Models in Low Precision — approx. low-precision training authors, 2024
https://scholar.google.com/scholar?q=The+Case+for+Training+Large+Models+in+Low+Precision15. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp316. AI Post Transformers: FlashFuser and Hopper-Era FFN Kernel Fusion — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-flashfuser-and-hopper-era-ffn-kernel-fus-e1fce9.mp317. AI Post Transformers: Automating DNN Compilation for FPGA Accelerators — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-automating-dnn-compilation-for-fpga-acce-6ef9bf.mp318. AI Post Transformers: Caffeine: A Unified FPGA for CNNs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-caffeine-a-unified-fpga-for-cnns-e8acbe.mp3