FlashAttention-4
Algorithm & Kernel Pipelining Co-Design for Asymmetric Hardware Scaling
Blackwell (B200) doubled tensor core throughput but left exp units and shared memory bandwidth in
the dust. FA4 redesigns everything to attack these new bottlenecks.