AI Post Transformers · Episode Companion

Adaptive Block-Scaled Data Types for FP4 Training

A new IF4 format from MIT and NVIDIA quantizes each 16-value group two ways — as FP4 and as scaled INT4 — and keeps whichever has lower error, encoding the choice for free in an otherwise-unused sign bit of the scale factor.

arXiv:2603.28765 Jack Cook et al. · MIT & NVIDIA · Mar 30, 2026 W4A4G4 Training + PTQ

Lineage: from FP8 to IF4

Each generation of 4-bit format traded away either representable values or dynamic range to control quantization error — until IF4 spent a bit that was already dead weight instead.

Read the diagram: FP8 (2022) set the block-scale template used at scale by DeepSeek-V3. NVFP4 and MXFP4 shrink to 4 bits with per-group scaling. 4/6 (Jan 2026) fixes NVFP4's error concentration by rescaling — at a cost. IF4 (Mar 2026) fixes the same problem for free.

Why 4 Bits at All: Matmul Speed on NVIDIA B200

4-bit matters primarily for raw multiply throughput, not memory — but only when both matmul operands are quantized, a bigger ask than compressing a finished checkpoint.

Relative matmul throughput, same hardware, both operands quantized to the listed format.

One Group, Two Candidate Encodings

For every 16-value group, IF4 quantizes twice — once as FP4, once as INT4 scaled by 7/6 — and keeps whichever has lower mean-squared error. Toggle between two example groups to see the winner flip.

The Free Bit: Scale Factor Sign

Scale factors are always positive in NVFP4 — their sign bit is dead weight. IF4 repurposes it as the FP4/INT4 indicator at zero storage cost.

8-bit E4M3 scale factor for one group of 16 values. Bit 7 (sign) is normally unused since scales are non-negative.

It Generalizes: IF3 and IF6

The same adaptive-choice pattern holds across bit widths — at a fixed memory budget, letting each group pick its own representation beats one fixed format for the whole tensor.

Training Loss vs. BF16 Baseline

340M-parameter dense transformer, ~100B FineWeb-Edu tokens, full W4A4G4 (weights, activations, gradients quantized in every linear layer except the last four hidden layers).

Final Relative Gap vs. BF16

Loss gap at the last checkpoint, same two recipes.

Where the Gain Actually Comes From

Ablation (Fig. 5a): hard-coding plain INT4 for the Hadamard-transformed weight gradient alone captures most of IF4's training benefit.

Honest caveat: within the training result alone, a static "use INT4 here" rule gets close to IF4's number. The adaptive machinery earns its keep more clearly in the PTQ setting, where the preprocessing regime isn't known in advance.

Perplexity: WikiText-2 vs. C4

Lower is better. Nemotron 3 Nano and Qwen3.5 checkpoints, up to the 397B-parameter MoE.

5-Task Accuracy — Qwen3.5 397B-MoE (17B active)

ARC-Easy, ARC-Challenge, HellaSwag, LAMBADA, PIQA. Dashed line marks the BF16 ceiling.

Average: IF4 72.8 vs. NVFP4 72.4, 4/6 72.3, MXFP4 70.5, against a 73.5 BF16 ceiling — lowest-or-tied-lowest perplexity in nearly every cell of Table 2.

MAC Unit: Decoding Two Formats from One Sign Bit

SystemVerilog MAC unit, synthesized in 28nm CMOS. Sixteen 4-bit weights and activations plus scale factors flow through a sign-bit check that routes each operand pair down one of two decode paths.

Isolated MAC Cost: NVFP4 Baseline vs. IF4

A bare MAC datapath is the worst-case showcase for IF4 — real processing elements spend most area/power on register files, buffering, and data movement, not multiplying.

Context: Recasens et al. ("Mind the Memory Gap," IEEE CLOUD 2025) make the same point about GPU inference generally — a few percent on one pipeline stage barely registers once you're bottlenecked on HBM bandwidth and data movement.

References

  1. Adaptive Block-Scaled Data Types for FP4 Training source paper
    Jack Cook, Hyemin Lee, Kathryn Le, Junxian Guo, Giovanni Traverso, Anantha Chandrakasan, Song Han · MIT / NVIDIA · 2026 · arxiv.org/abs/2603.28765
  2. Mixed Precision Training
    Micikevicius, Narang, Alben, Diamos, Elsen, Garcia, Ginsburg, Houston, Kuchaiev, Venkatesh, Wu · NVIDIA/Baidu · ICLR 2018 · scholar link
  3. FP8 Formats for Deep Learning
    Micikevicius, Stosic, Burgess, Cornea, Dubey, Grisenthwaite, Ha, Heinecke, Judd, Kamalu, Mellempudi, Oberman, Shoeybi, Siu, Wu · NVIDIA, Arm, Intel, Qualcomm · 2022 · scholar link
  4. Microscaling Data Formats for Deep Learning
    Rouhani, Zhao, More, et al. · Microsoft, AMD, Arm, Intel, Meta, NVIDIA, Qualcomm (OCP) · 2023 · scholar link
  5. DeepSeek-V3 Technical Report
    DeepSeek-AI · 2024 · scholar link
  6. Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling
    Jack Cook, Junxian Guo, Guangxuan Xiao, Yujun Lin, Song Han · 2026 · scholar link
  7. Pretraining Large Language Models with NVFP4
    NVIDIA · 2026 · scholar link
  8. Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
    Andrei Panferov, Erik Schultheis, Soroush Tabesh, Dan Alistarh · IST Austria · 2026 · scholar link
  9. INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
    Mengzhao Chen, Meng Wu, Hui Jin, Zhihang Yuan, et al. · 2025 · scholar link
  10. Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
    Vage Egiazarian, Roberto L. Castro, Denis Kuznedelev, et al. · 2026 · scholar link
  11. WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
    Jiale Chen, Vage Egiazarian, Roberto L. Castro, Torsten Hoefler, Dan Alistarh · 2026 · scholar link
  12. Scaling Laws for Precision
    Tanishq Kumar, Zachary Ankner, Benjamin F. Spector, et al. · 2024 · scholar link