AI Post Transformers Companion arXiv: 2309.14592 MLSys 2024 Transcript IDs: loading

Efficient Post-Training Quantization with FP8

FP8 changes the numeric bargain after training: more dynamic range than plain INT8, enough precision to preserve most transformer behavior, and fewer outlier-induced fallbacks. This page turns the episode into a visual lab for format choice, operator placement, BatchNorm recalibration, and the headline coverage gap across 75 architectures and 200+ task cases.

Primary Result
FP8 coverage 92.64% vs INT8 65.87%
Format Story
E4M3 leads many NLP cases; E3M4 edges some vision
Operational Hook
Quantize what kernels support; keep critical accumulators higher precision
Important Caveat
Numerical viability shown first; native FP8 system speed still depends on hardware
Pipeline Diagram

PTQ Recipe, Not Magic Datatype

The paper's core contribution is procedural: calibrate, place formats by operator, extend quantization only where kernels and accuracy allow, then repair normalization drift before deployment.

Quantized path Format choice / tuning Higher-precision anchor Fallback pressure
Bargain Map

Range vs Precision vs Coverage

Click a format point. The best deployment format is not the widest or the most precise in isolation; it is the one that lands in the right workload region with the fewest unsupported operators and the least outlier damage.

Select a format point to inspect its numeric tradeoff.
92.64%
Reported overall workload coverage for FP8 PTQ.
65.87%
Reported overall workload coverage for the INT8 comparison baseline.
75
Architectures spanning NLP, vision, segmentation, and image generation.
200+
Task cases used to test whether quantized inference survives accuracy thresholds.
96.32%
NLP coverage reported for E4M3, the balanced FP8 format.
78.95%
Vision coverage reported for E3M4, which slightly beats E4M3 there.
Interactive Format Lab

Bit Layout and Representable Span

Switch among INT8 and three FP8 variants. The exponent buys range, the mantissa buys local detail, and the winning format depends on whether your model fails from overflow or from fine-grained rounding.

Heatmap

Quantization Stress Across Magnitudes

Each cell is a mock error surface across activation magnitude and calibration aggressiveness. FP8 makes the cliff less abrupt because the exponent spreads representable values over a wider dynamic range.

Hover a cell to inspect relative clipping and rounding stress.
Comparison Modes

Coverage, Accuracy Retention, and Fallback Rate

The paper's cleanest headline is coverage, not raw latency. This view separates "how often it works" from "how much score it keeps" and from "how often the graph crawls back to higher precision."

Switch modes or hover a bar to inspect benchmark slices.
Workload Split

E4M3 vs E3M4 by Domain

The episode's useful nuance is that FP8 is not one format. The balanced middle format looks strongest for language, while the more precision-tilted E3M4 occasionally wins on vision-like workloads.

E4M3 E3M4 E5M2 as range insurance
Placement Diagram

Which Operators Stay Quantized?

This is where hardware-awareness cashes out. Big matrix multiplies are the easy residents; normalization, softmax, residual edges, and accumulators decide whether a recipe is deployable instead of numerically pretty on paper.

Residency Heatmap

Low-Precision Reach by Model Family

The extended recipe improves operator reach, especially where BatchNorm recalibration or broader kernel support lets more of the graph stay in low precision.

Hover a heatmap cell to inspect low-precision operator residency.
Activation Matrix

Outlier Damage: INT8 vs FP8

This heatmap shows a synthetic token-by-channel activation slice. INT8 suffers once a few channels stretch the global scale, while FP8 absorbs more of the spike with exponent headroom before the rest of the matrix is sacrificed.

Hover a cell to inspect magnitude, clipping pressure, and relative error.
Scale Geometry

Uniform Bins vs Exponent Ladder

The top chart tracks error as activation percentiles get nastier. The lower band sketches the representable geometry: equal INT8 bins with an external scale versus expanding FP8 steps that follow magnitude.

References

Compact Source Shelf

Explicit transcript/source arXiv extraction yields one ID: 2309.14592. Additional references below follow the source list provided for the episode.

01
Efficient Post-training Quantization with FP8 Formats Haihao Shen, Naveen Mellempudi, Xin He, Qun Gao, Chang Wang, Mengni Wang. Primary paper discussed in the episode.
arXiv
02
Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference Canonical INT8 deployment reference cited in the discussion.
Scholar
03
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale Important transformer-specific INT8 baseline for structured outlier handling.
arXiv
04
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models Referenced in the episode as the stronger INT8-style language baseline family.
arXiv
05
FP8 Formats for Deep Learning Modern FP8 format framing for E4M3 and E5M2.
Scholar
06
FP8 Quantization: The Power of the Exponent Why exponent bits help on activation outliers and dynamic range.
Scholar
07
AI Post Transformers: Deep Kernel Fusion for Transformer Decoding Related episode cited in the transcript for end-to-end decode bottlenecks.
Audio
08
AI Post Transformers: Mooncake for KV Cache-Centric LLM Serving Related episode cited in the transcript for memory-bound serving constraints.
Audio