← All episodes Adaptive Block-Scaled Data Types for FP4 Training

Adaptive Block-Scaled Data Types for FP4 Training

Jul 30, 2026
This episode explores Adaptive Block-Scaled Data Types, a new IF4 format from MIT and NVIDIA researchers for representing numbers in just 4 bits during LLM training and inference. The discussion traces the lineage from FP8 training (used at scale by DeepSeek-V3) through existing 4-bit formats like NVFP4 and MXFP4, and the predecessor "4/6" method, explaining why each prior approach traded away either representable values or dynamic range to control quantization error. The key innovation covered is how IF4 quantizes each 16-value group both as FP4 and as scaled INT4, keeping whichever has lower error, and encodes that choice for free in an otherwise-unused sign bit of the scale factor. Listeners get a clear picture of why 4-bit precision matters primarily for raw matmul speed on hardware like NVIDIA's B200, not just memory savings, and why this fix is notable for spending "dead weight" bits rather than sacrificing precision or range like earlier techniques.
Sources:
1. Adaptive Block-Scaled Data Types for FP4 Training
https://arxiv.org/pdf/2603.28765
2. Mixed Precision Training — Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, Hao Wu (NVIDIA/Baidu), 2017/2018 (ICLR 2018)
https://scholar.google.com/scholar?q=Mixed+Precision+Training
3. FP8 Formats for Deep Learning — Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep Dubey, Richard Grisenthwaite, Sangwon Ha, Alexander Heinecke, Patrick Judd, John Kamalu, Naveen Mellempudi, Stuart Oberman, Mohammad Shoeybi, Michael Siu, Hao Wu (NVIDIA, Arm, Intel, Qualcomm), 2022
https://scholar.google.com/scholar?q=FP8+Formats+for+Deep+Learning
4. Microscaling Data Formats for Deep Learning — Bita Darvish Rouhani, Ritchie Zhao, Ankit More, Mathew Hall, Alireza Khodamoradi, Summer Deng, Dhruv Choudhary, Marius Cornea, Eric Dellinger, Kristof Denolf, Stosic Dusan, Venmugil Elango, Maximilian Golub, Alexander Heinecke, Phil James-Roxby, Dharmesh Jani, Gaurav Kolhe, Martin Langhammer, Ada Li, Levi Melnick, Maral Mesmakhosroshahi, Andres Rodriguez, Michael Schulte, Rasoul Shafipour, Lei Shao, Michael Siu, Pradeep Dubey, Paulius Micikevicius (Microsoft, AMD, Arm, Intel, Meta, NVIDIA, Qualcomm — OCP consortium), 2023
https://scholar.google.com/scholar?q=Microscaling+Data+Formats+for+Deep+Learning
5. DeepSeek-V3 Technical Report — DeepSeek-AI (large author list, DeepSeek), 2024
https://scholar.google.com/scholar?q=DeepSeek-V3+Technical+Report
6. Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling — Jack Cook, Junxian Guo, Guangxuan Xiao, Yujun Lin, Song Han, 2026
https://scholar.google.com/scholar?q=Four+Over+Six%3A+More+Accurate+NVFP4+Quantization+with+Adaptive+Block+Scaling
7. Pretraining Large Language Models with NVFP4 — NVIDIA (large author list), 2026
https://scholar.google.com/scholar?q=Pretraining+Large+Language+Models+with+NVFP4
8. Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation — Andrei Panferov, Erik Schultheis, Soroush Tabesh, Dan Alistarh, 2026
https://scholar.google.com/scholar?q=Quartet+II%3A+Accurate+LLM+Pre-Training+in+NVFP4+by+Improved+Unbiased+Gradient+Estimation
9. INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats — Mengzhao Chen, Meng Wu, Hui Jin, Zhihang Yuan, et al., 2025
https://scholar.google.com/scholar?q=INT+v.s.+FP%3A+A+Comprehensive+Study+of+Fine-Grained+Low-bit+Quantization+Formats
10. Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization — Vage Egiazarian, Roberto L. Castro, Denis Kuznedelev, et al., 2026
https://scholar.google.com/scholar?q=Bridging+the+Gap+Between+Promise+and+Performance+for+Microscaling+FP4+Quantization
11. WUSH: Near-Optimal Adaptive Transforms for LLM Quantization — Jiale Chen, Vage Egiazarian, Roberto L. Castro, Torsten Hoefler, Dan Alistarh, 2026
https://scholar.google.com/scholar?q=WUSH%3A+Near-Optimal+Adaptive+Transforms+for+LLM+Quantization
12. Scaling Laws for Precision — Tanishq Kumar, Zachary Ankner, Benjamin F. Spector, et al., 2024
https://scholar.google.com/scholar?q=Scaling+Laws+for+Precision
Interactive Visualization: Adaptive Block-Scaled Data Types for FP4 Training