FP8 changes the numeric bargain after training: more dynamic range than plain INT8, enough precision to preserve most transformer behavior, and fewer outlier-induced fallbacks. This page turns the episode into a visual lab for format choice, operator placement, BatchNorm recalibration, and the headline coverage gap across 75 architectures and 200+ task cases.
The paper's core contribution is procedural: calibrate, place formats by operator, extend quantization only where kernels and accuracy allow, then repair normalization drift before deployment.
Click a format point. The best deployment format is not the widest or the most precise in isolation; it is the one that lands in the right workload region with the fewest unsupported operators and the least outlier damage.
Switch among INT8 and three FP8 variants. The exponent buys range, the mantissa buys local detail, and the winning format depends on whether your model fails from overflow or from fine-grained rounding.
Each cell is a mock error surface across activation magnitude and calibration aggressiveness. FP8 makes the cliff less abrupt because the exponent spreads representable values over a wider dynamic range.
The paper's cleanest headline is coverage, not raw latency. This view separates "how often it works" from "how much score it keeps" and from "how often the graph crawls back to higher precision."
The episode's useful nuance is that FP8 is not one format. The balanced middle format looks strongest for language, while the more precision-tilted E3M4 occasionally wins on vision-like workloads.
This is where hardware-awareness cashes out. Big matrix multiplies are the easy residents; normalization, softmax, residual edges, and accumulators decide whether a recipe is deployable instead of numerically pretty on paper.
The extended recipe improves operator reach, especially where BatchNorm recalibration or broader kernel support lets more of the graph stay in low precision.
This heatmap shows a synthetic token-by-channel activation slice. INT8 suffers once a few channels stretch the global scale, while FP8 absorbs more of the spike with exponent headroom before the rest of the matrix is sacrificed.
The top chart tracks error as activation percentiles get nastier. The lower band sketches the representable geometry: equal INT8 bins with an external scale versus expanding FP8 steps that follow magnitude.
Explicit transcript/source arXiv extraction yields one ID: 2309.14592. Additional references below follow the source list provided for the episode.