AI Post Transformers / Visual Companion

SageAttention2 and Fast Exact INT4 Attention

SageAttention2 is not a new attention rule. It is a hardware-software argument that exact dense attention can move deeper into INT4 and FP8 only if the numerical damage is repaired at the same granularity that GPU threads already use to execute mma fragments.

ICML 2025 lineage Exact dense attention Q,K -> INT4 P~ , V -> FP8 Transcript scan: no extra arXiv IDs beyond 2411.10958 arXiv:2411.10958

Visual Thesis

Low-bit exact attention works only when precision, scaling, and kernel layout are co-designed instead of treated as separate optimizations.

Headline kernel ~481 TOPS on RTX 4090
Application uplift ~1.7x to 1.8x
Numerical patch Q correction + foldback
Constraint Head dim padded to 64/128

The Quadratic Bill

Exact kernels matter because they preserve behavior, but long context still drags them up the same rising curve. The paper's promise is not to escape O(n^2), only to make that curve cheaper.

exact path repaired low precision changed semantics

Relative Attention Time vs Sequence Length

Illustrative wall-time scaling: exact low-bit kernels reduce the constant factor, not the asymptotic shape.

Attention Map Fidelity Under Quantization

Hover cells to inspect how naive INT4 can shift softmax winners and how repaired INT4 pulls the map back.

Numerical Repair Stack

The page's central claim is visual here: low bits alone do not help. The system works only because smoothing, exact INT4 fragments, FP8 value mixing, and periodic FP32 foldback are chained together.

Step-by-Step Exact Attention Pipeline

Click a stage. The highlighted block shows where the paper spends engineering effort to keep the exact rule intact.

Channel Outlier Smoothing

Paired bars show how a few dominant channels stop monopolizing the quantization scale budget.
Before Scale hijack Three channels dominate range, leaving ordinary values too coarse for INT4.
After Range redistribution Channelwise repair flattens the tail enough to preserve logit ordering.
Why it matters Softmax is the amplifier Small QK errors become routing errors once the exponentials decide who wins.

Per-Thread INT4 Geometry

Per-thread quantization is the systems move that makes the paper distinct. Scale granularity gets finer, but the bookkeeping stays attached to thread-owned fragments that already flow into PTX mma execution.

Fragment Map Aligned to Thread Ownership

Toggle coarse per-block scaling against the paper's per-thread layout.

Granularity Tradeoff

Accuracy risk falls with finer scales. The paper's claim is that thread alignment keeps overhead from erasing the win.
Implementation envelope GQA and varlen exposed The public code path supports different q and kv lengths but pads head dim to 64 or 128 and rejects larger than 128.
Skeptical read Not effortless portability The method is strongest in selected NVIDIA inference regimes, not as a universal drop-in proof.

Benchmark Accounting

The flashy number is a kernel number. The more interesting chart is the second one: how much of that win survives quantization, smoothing, KV plumbing, and serving-framework overhead.

Compare Deployment Contexts

Switch GPU context to see where the transcript cites measured results and where the comparison becomes more about framing.

Kernel Win vs Serving Win

Illustrative waterfall based on the transcript's boundary between kernel-only throughput and end-to-end application speedup.
Headline About 3x kernel speedup RTX 4090 figures are shown from the chart's good side: kernel only, without full preprocessing overhead.
Sober takeaway About 1.7x to 1.8x realized That smaller number is still real, and it is the one that matters for dense long-context deployments.

References

Compact arXiv-linked lineage: exact kernels, low-bit attention, outlier handling, and sparse long-context alternatives.