arXiv:2603.05451 — Zadouri et al.

FlashAttention-4 Algorithm & Kernel Pipelining Co-Design for Asymmetric Hardware Scaling

Blackwell (B200) doubled tensor core throughput but left exp units and shared memory bandwidth in the dust. FA4 redesigns everything to attack these new bottlenecks.