Lineage: from FP8 to IF4
Each generation of 4-bit format traded away either representable values or dynamic range to control quantization error — until IF4 spent a bit that was already dead weight instead.
Why 4 Bits at All: Matmul Speed on NVIDIA B200
4-bit matters primarily for raw multiply throughput, not memory — but only when both matmul operands are quantized, a bigger ask than compressing a finished checkpoint.
One Group, Two Candidate Encodings
For every 16-value group, IF4 quantizes twice — once as FP4, once as INT4 scaled by 7/6 — and keeps whichever has lower mean-squared error. Toggle between two example groups to see the winner flip.
The Free Bit: Scale Factor Sign
Scale factors are always positive in NVFP4 — their sign bit is dead weight. IF4 repurposes it as the FP4/INT4 indicator at zero storage cost.
It Generalizes: IF3 and IF6
The same adaptive-choice pattern holds across bit widths — at a fixed memory budget, letting each group pick its own representation beats one fixed format for the whole tensor.
Training Loss vs. BF16 Baseline
340M-parameter dense transformer, ~100B FineWeb-Edu tokens, full W4A4G4 (weights, activations, gradients quantized in every linear layer except the last four hidden layers).
Final Relative Gap vs. BF16
Loss gap at the last checkpoint, same two recipes.
Where the Gain Actually Comes From
Ablation (Fig. 5a): hard-coding plain INT4 for the Hadamard-transformed weight gradient alone captures most of IF4's training benefit.
Perplexity: WikiText-2 vs. C4
Lower is better. Nemotron 3 Nano and Qwen3.5 checkpoints, up to the 397B-parameter MoE.
5-Task Accuracy — Qwen3.5 397B-MoE (17B active)
ARC-Easy, ARC-Challenge, HellaSwag, LAMBADA, PIQA. Dashed line marks the BF16 ceiling.
MAC Unit: Decoding Two Formats from One Sign Bit
SystemVerilog MAC unit, synthesized in 28nm CMOS. Sixteen 4-bit weights and activations plus scale factors flow through a sign-bit check that routes each operand pair down one of two decode paths.
Isolated MAC Cost: NVFP4 Baseline vs. IF4
A bare MAC datapath is the worst-case showcase for IF4 — real processing elements spend most area/power on register files, buffering, and data movement, not multiplying.
References
- Adaptive Block-Scaled Data Types for FP4 Training source paper
- Mixed Precision Training
- FP8 Formats for Deep Learning
- Microscaling Data Formats for Deep Learning
- DeepSeek-V3 Technical Report
- Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling
- Pretraining Large Language Models with NVFP4
- Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
- INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
- Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
- WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
- Scaling Laws for Precision