SageAttention2 is not a new attention rule. It is a hardware-software argument that exact dense attention can move deeper into INT4 and FP8 only if the numerical damage is repaired at the same granularity that GPU threads already use to execute mma fragments.
Low-bit exact attention works only when precision, scaling, and kernel layout are co-designed instead of treated as separate optimizations.
Exact kernels matter because they preserve behavior, but long context still drags them up the same rising curve. The paper's promise is not to escape O(n^2), only to make that curve cheaper.
The page's central claim is visual here: low bits alone do not help. The system works only because smoothing, exact INT4 fragments, FP8 value mixing, and periodic FP32 foldback are chained together.
Per-thread quantization is the systems move that makes the paper distinct. Scale granularity gets finer, but the bookkeeping stays attached to thread-owned fragments that already flow into PTX mma execution.
The flashy number is a kernel number. The more interesting chart is the second one: how much of that win survives quantization, smoothing, KV plumbing, and serving-framework overhead.
Compact arXiv-linked lineage: exact kernels, low-bit attention, outlier handling, and sparse long-context alternatives.