1 00:00:01,000 --> 00:00:36,024 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today's paper is Adaptive Block-Scaled Data Types, from Jack Cook et al. — seven authors in all, including Hyemin Lee, Kathryn Le, Junxian Guo, Giovanni Traverso, Anantha Chandrakasan, and Song Han — out of MIT and NVIDIA, posted to arXiv on March 30th, 2026. It's about a new way to pack numbers into 4 bits that's supposed to beat what everyone's currently using for training and running language models at that precision. 2 00:00:36,024 --> 00:00:54,650 [Dr. Ada Shannon] What pulled me in is that this isn't a format someone hacked together and threw a benchmark at — the whole design falls out of one clean observation about exactly where existing 4-bit formats bleed accuracy. They diagnose the problem precisely, and the fix follows almost mechanically from that diagnosis. 3 00:00:54,650 --> 00:01:11,975 [Hal Turing] Before we get into what they built, let's set the stakes for people who think "quantization" just means shrinking a model to fit on a laptop. Why does anyone care about 4-bit numbers during training itself — is the payoff here really about raw speed, not just memory savings? 4 00:01:11,975 --> 00:01:56,600 [Dr. Ada Shannon] Speed, mostly. On NVIDIA's B200 or AMD's MI355X, FP4 matmuls run roughly twice as fast as FP8, and about four times as fast as FP16 or BF16 — but only if both operands of the multiply are quantized, a bigger ask than shrinking a finished checkpoint. The catch: a 4-bit float can only represent 16 values total, so you can't just chop a tensor down directly. The fix is block scaling — split the tensor into small groups and give each its own scale factor. NVFP4 uses groups of 16 sharing an FP8 scale factor; MXFP4 uses groups of 32 sharing a coarser power-of-two scale, standardized by Rouhani and a big multi-company group spanning Microsoft, AMD, Meta and NVIDIA in the 2023 OCP microscaling spec. 5 00:01:56,600 --> 00:02:16,900 [Hal Turing] Got it — each group carries its own local range, and the scale factors carry the big picture. One more bit of notation, since it'll matter later — what's the actual difference between quantizing a finished model versus quantizing it while it's training? And what does the extra G in W4A4G4 mean? 6 00:02:16,900 --> 00:02:37,350 [Dr. Ada Shannon] W4A4 means weights and activations are both quantized to 4 bits after training is already done — that's PTQ, the world of GPTQ and AWQ, compressing a finished checkpoint. W4A4G4 adds the G for gradients, quantizing weights, activations, AND gradients to 4 bits during the training loop itself, which is much harder because— 7 00:02:37,350 --> 00:02:47,450 [Hal Turing] Oh wait wait wait — because the errors don't just show up once, they compound, every single step feeding a slightly wrong gradient into the next update. 8 00:02:47,450 --> 00:03:28,750 [Dr. Ada Shannon] Exactly. So FP4 training recipes lean on extra machinery: the random Hadamard transform, a fixed rotation applied before quantization that smooths outliers so they don't blow up a group's scale factor, and stochastic rounding, which rounds up or down probabilistically instead of always to nearest, keeping gradient estimates unbiased across millions of steps. None of it's free. This lineage traces to Micikevicius and colleagues at NVIDIA, whose FP8 Formats for Deep Learning paper in 2022 defined the scale format this family uses — and it's proven at real scale, too: DeepSeek trained their 671-billion-parameter DeepSeek-V3 largely in FP8 in 2024 the same way. 9 00:03:28,750 --> 00:03:38,725 [Hal Turing] So that's the general survival kit for FP4 training. Where does this paper's direct predecessor, Four Over Six, fit into that picture? 10 00:03:38,725 --> 00:04:12,500 [Dr. Ada Shannon] Four Over Six — 4/6 — is Jack Cook again, with Junxian Guo, Guangxuan Xiao, Yujun Lin and Song Han, MIT, January 2026. NVFP4 always scales a group so its largest value hits 6, concentrating error on near-maximal values. 4/6 lets a group instead scale its max to 4, spreading error more evenly. It works, but costs two things: giving up two of your fifteen usable FP4 values, and shrinking the whole tensor's global scale to stop overflow, cutting dynamic range by nearly 43 percent. 11 00:04:12,500 --> 00:04:32,975 [Hal Turing] I want to push back here, Ada. You're describing a format that already needs a Hadamard transform and stochastic rounding, and now an adaptive rescaling trick that costs values and dynamic range on top. Doesn't that just mean 4-bit isn't ready, and you should stay at FP8, which you just told me works at DeepSeek scale? 12 00:04:32,975 --> 00:04:51,775 [Dr. Ada Shannon] No, I don't think that follows. Every one of those tricks is nearly free at inference — you're rearranging bits you already have, not adding a slow kernel. And 4/6 isn't "FP4 doesn't work," it's "this specific fix has a specific, measurable bill." That's exactly why they went looking for the same idea without paying it. 13 00:04:51,775 --> 00:05:02,850 [Hal Turing] Okay — if the fix spends a bit that was already dead weight instead of a value you needed, that's a genuinely different kind of trade than what 4/6 made. I'll take that. 14 00:05:02,850 --> 00:05:24,125 [Dr. Ada Shannon] That's exactly the pitch. For every group of 16 values, quantize it two ways — once as FP4, once as scaled INT4 — and keep whichever has lower error. The choice gets flagged for free in the sign bit of the group's scale factor, which is normally dead weight in NVFP4, since scale factors are always positive. Zero extra storage, one adaptive decision per group. They call it IF4. 15 00:05:24,125 --> 00:05:36,325 [Hal Turing] Wait, they actually pronounce that "eye-eff-four"? Somewhere down the line some engineer is going to write "if IF4" in their kernel code and find out the hard way whether their linter survives it. 16 00:05:36,325 --> 00:05:50,175 [Dr. Ada Shannon] That linter's having a rough day either way. That's the headline — the actual bookkeeping behind sharing dynamic range between the two representations without 4/6's problems is worth walking through properly. 17 00:05:50,175 --> 00:06:21,850 [Dr. Ada Shannon] Here's the bookkeeping. For each group of sixteen values, IF4 computes the block scale like NVFP4 does — largest value over six, cast to FP8 E4M3. Then it quantizes twice, once as FP4, once as INT4 scaled by seven-sixths, and keeps whichever gives lower mean squared error. The choice is stored for free: every value already has its own sign bit, so the group's shared scale factor never needed one. IF4 repurposes that unused bit as the indicator — zero for FP4, one for INT4 — with zero memory overhead. 18 00:06:21,850 --> 00:06:39,325 [Hal Turing] Clever — instead of losing two representable values the way Four Over Six did by capping at four, you're choosing between two full fifteen-value alphabets and letting the data decide. Does that trick only work at four bits, or does it scale to other precisions? 19 00:06:39,325 --> 00:07:01,250 [Dr. Ada Shannon] They generalize it. IF3 makes the same choice at three bits, between a tiny E2M0 float and a scaled integer. IF6 goes to six bits with two floating-point flavors, E2M3 and E3M2, each with its own rescaling constant. The pattern in Figure 2 holds across the family — at a fixed memory budget, letting each group pick its representation beats one fixed format for the whole tensor. 20 00:07:01,250 --> 00:07:15,076 [Hal Turing] It's a whole family, not a one-off hack. So what's the receipt — how did they actually stress-test this in a real training run, where errors compound over thousands of steps instead of sitting in a static error table? 21 00:07:15,076 --> 00:07:52,001 [Dr. Ada Shannon] They pre-trained a 340-million-parameter dense transformer with query-key normalization on about a hundred billion tokens of FineWeb-Edu, in a recipe close to NVIDIA's own NVFP4 paper. Full W4A4G4 — weights, activations, and gradients quantized in every linear layer except the final four hidden layers, random Hadamard transform on the weight gradient, stochastic rounding on activation gradients. Same survival kit as before, just with IF4 in for NVFP4. Across that run, IF4's training loss tracks measurably closer to the BF16 baseline than NVFP4 does. 22 00:07:52,001 --> 00:07:59,776 [Hal Turing] And you mentioned there's a setup where that gap gets even wider — something about pairing it with another gradient trick? 23 00:07:59,776 --> 00:08:20,526 [Dr. Ada Shannon] MS-EDEN, from Panferov, Schultheis, Tabesh, and Alistarh's Quartet II paper, IST Austria, January 2026. It Hadamard-transforms every backward-pass input, not just the weight gradient. Pair it with IF4 and the relative-loss gap versus BF16 in Figure 4 widens further — IF4 pulls ahead exactly when more of the backward pass gets smoothed toward uniform. 24 00:08:20,526 --> 00:08:35,676 [Hal Turing] I'll push back there, Ada. Isn't that a flattering setup? You're crediting IF4 for a gap that only really opens up once you've bolted on someone else's gradient estimator — feels like two papers sharing credit for one number. 25 00:08:35,676 --> 00:08:55,351 [Dr. Ada Shannon] I actually disagree with you there, Hal. Look at the standalone number — no MS-EDEN, just the standard recipe — IF4 is already ahead of NVFP4. MS-EDEN doesn't manufacture that gap, it amplifies one that already exists, because transforming more of the backward pass pushes more groups toward the distribution IF4's INT4 branch handles well. 26 00:08:55,351 --> 00:09:06,226 [Hal Turing] Fair — the baseline number's the real story, MS-EDEN just makes it more visible. Let's hit the other half of the paper — what did the PTQ results look like? 27 00:09:06,226 --> 00:09:52,926 [Dr. Ada Shannon] Two ways. Table 2 reports WikiText-2 and C4 perplexity across NVIDIA's Nemotron 3 Nano and Alibaba's Qwen3.5, up through their 122-billion-parameter MoE checkpoints, and IF4 is lowest or tied-lowest in nearly every cell. Table 3 averages accuracy across five tasks — ARC-Easy, ARC-Challenge, HellaSwag, LAMBADA, PIQA — scaling to Qwen3.5's 397-billion-parameter, 17-billion-active MoE. IF4 lands at 72.8 average against a 73.5 BF16 ceiling, ahead of NVFP4's 72.4 and Four Over Six's 72.3, clear of MXFP4's 70.5 — winning the simpler PTQ setting too, at real production sizes. 28 00:09:52,926 --> 00:10:04,776 [Hal Turing] All of that's simulated in software, though. Did they check whether decoding two number formats out of one sign bit is cheap in actual silicon, or is that left as future work? 29 00:10:04,776 --> 00:10:27,002 [Dr. Ada Shannon] No hand-wave — they built a MAC unit in SystemVerilog and synthesized it in 28-nanometer CMOS. It takes sixteen 4-bit weights and activations plus scale factors, checks the sign bit on each, and routes NVFP4 operands through a lookup table while INT4 operands go through shifter logic. Products land in FP16, the scale factors multiply in a separate path, and depending on which format combination showed up— 30 00:10:27,002 --> 00:10:32,702 [Hal Turing] Hold on — that's a lot of extra plumbing per multiply. That has to cost something. 31 00:10:32,702 --> 00:11:11,952 [Dr. Ada Shannon] It does, just not much. A mixed FP4-and-INT4 pair needs a six-sevenths range-alignment factor; two INT4 operands need thirty-six forty-ninths — one extra FP32 multiply before accumulation. Against a plain NVFP4 baseline, the full unit synthesizes to 4.7% higher latency. But that's an isolated MAC number — real throughput on these accelerators is usually capped by HBM bandwidth and data movement. That's the same point Recasens and colleagues made in 'Mind the Memory Gap' at IEEE CLOUD 2025 about GPU inference generally — a few percent on one pipeline stage barely registers once you're bottlenecked on moving bytes. 32 00:11:11,952 --> 00:11:34,402 [Hal Turing] So that 4.7 percent latency number was just one MAC unit in isolation, not a full accelerator. Once you zoom out to what a chip designer actually cares about — area and power — what does adopting IF4 really cost? Because "slightly slower per multiply" and "needs more silicon per multiply" are two very different pitches to an ASIC team. 33 00:11:34,402 --> 00:12:39,377 [Dr. Ada Shannon] They're blunter than the latency number suggests. In their 28-nanometer synthesis, the IF4 MAC uses 66.6 percent more area and draws 27.8 percent more power than a plain NVFP4 baseline. NVFP4 also clocked higher — 524 megahertz with slack to spare, versus IF4 meeting timing at exactly 500 megahertz with zero slack, a 4.6 percent throughput hit on top of the latency gap. Two things drive it: decoded operands need a slightly wider fixed-point format to hold both FP4 and INT4 magnitudes, and there's an entirely new FP32 multiplier for range alignment that a single-format MAC never needs. The paper's counter-argument is that a bare MAC datapath is the worst possible showcase — a real processing element spends most of its area on register files and buffering, and most of its power moving data to and from memory, not multiplying. 34 00:12:39,377 --> 00:13:07,003 [Hal Turing] Fair, but let's push on the bigger claim in this paper, because something's been bothering me. Every actual training number comes from one 340-million-parameter dense transformer, one hyperparameter configuration, a hundred billion FineWeb-Edu tokens. The PTQ side, meanwhile, scales all the way up to Qwen3.5 at 397 billion parameters. Those are not the same weight class of evidence for a paper whose headline claim is about training. 35 00:13:07,003 --> 00:13:50,728 [Dr. Ada Shannon] They're really not, and it matters because PTQ is the easier problem — you quantize a finished model once and measure the damage. Training compounds: a slightly biased gradient at step one thousand feeds a slightly wrong weight into step one thousand and one, for hundreds of thousands of steps, and that's exactly the regime never tested here. Worth noting — this paper and its baseline, Four Over Six, share the same lead authors and MIT lab: Jack Cook, Junxian Guo, and Song Han. No outside group has replicated either result yet. And at the scale they did test, on Qwen3.5 at 397 billion parameters, IF4 averages 80.94 versus NVFP4's 80.90 — that's a three-run average with no reported variance, which is asking a lot of a 0.04-point gap. 36 00:13:50,728 --> 00:14:17,528 [Hal Turing] Wait, hold on — actually, that ties into something in Figure 5a that bugged me. You said the training gap traces mostly to IF4 preferring INT4 on Hadamard-transformed weight gradients. But their own ablation shows just hard-coding plain NVINT4 for that one computation, no adaptivity, no indicator bit, gets you almost all the way to IF4's number. Doesn't that mean the actual novel contribution here is smaller than the headline? 37 00:14:17,528 --> 00:14:44,828 [Dr. Ada Shannon] That's a fair read of that one figure, and I won't dress it up. A static rule — use INT4 for the Hadamard-transformed weight gradient, full stop — captures most of the training gain by itself. But look at the PTQ table instead: raw, untransformed weights favor NVFP4, Hadamard-transformed weights favor NVINT4, and IF4 comes out ahead in both columns because it doesn't need to know in advance which regime it's in. A hard-coded rule only works if you already know your pipeline's preprocessing. 38 00:14:44,828 --> 00:15:06,178 [Hal Turing] I actually disagree with you there, Ada. If the ablation shows a dumb static swap gets nearly the whole training benefit, the burden's on the adaptive machinery to prove it's earning its keep on the training claim specifically — "it also helps at inference" is answering a different question than the one the headline result is making. 39 00:15:06,178 --> 00:15:25,778 [Dr. Ada Shannon] No — I think you're merging two separate experiments into one verdict. I'm not using the PTQ result to rescue the training figure; they're independent claims standing on their own evidence. But you're right that within training alone, the paper undersells how much of that number is just "use INT4 here," and it should say so louder. I'll give you that much. 40 00:15:25,778 --> 00:15:43,453 [Hal Turing] I'll take it. So beyond that, there's the bias question from Appendix B — they admit IF4 inherits some quantization bias from Four Over Six under stochastic rounding, unlike vanilla NVFP4 which stays unbiased. How worried should we be? 41 00:15:43,453 --> 00:16:31,178 [Dr. Ada Shannon] Modestly, and they're honest about the caveat: it's measured on the same single 340-million-parameter run, and the paper says outright that larger-scale experiments will likely be needed to confirm it stays negligible. Practically, though, this format is a strong drop-in candidate — zero memory overhead versus NVFP4, the indicator bit is free real estate in a sign bit nobody was using, and it targets exactly the Blackwell and MI355X-class hardware already shipping NVFP4 support. The open question for future work is whether that edge survives contact with calibrated PTQ methods like MR-GPTQ or WUSH instead of the naive round-to-nearest used throughout this paper. 42 00:16:31,178 --> 00:17:03,753 [Hal Turing] So here's where I land: IF4 is a genuinely clever, low-cost idea — free bit, no memory tax, real gains in the PTQ regime that's tested at real scale. The training story is promising but thin, resting on one small run, and a chunk of its advantage traces back to a much simpler static rule than the adaptive framing suggests. Worth watching, not yet worth calling a settled win. That's Adaptive Block-Scaled Data Types and IF4 — thanks for listening, everyone, and we'll catch you next time.