STEEL fuses FlashAttention-2's tiling into a three-stage pipeline purpose-built for AMD's spatial-dataflow XDNA NPU, and fixes the load imbalance the causal mask creates on a fixed pipeline by scattering rows across pipelines instead of assigning contiguous chunks.
A GPU is SIMT with a cache hierarchy deciding what stays close to compute. XDNA is a 2-D grid of VLIW tiles wired by an on-chip network — the compiler, not a cache, decides which tile does what and schedules every transfer.
STEEL splits FlashAttention-2's math across three dedicated AIE cores connected by IRON ObjectFIFOs — typed, synchronized queues — instead of one core looping through every step serially.
Roughly half the attention matrix is masked to zero. Assigning contiguous row-blocks per pipeline gives early pipelines mostly-masked (cheap) tiles and late pipelines full unmasked (expensive) tiles — since K/V broadcast waits for every consumer, the slowest pipeline stalls the whole array.
Each Mem tile has only 6 ports. Every STEEL pipeline needs 4: one distributed Q tile, one collected O tile, and two for swizzling intermediate P tiles between the softmax and PV stages. K/V are broadcast, shared across all pipelines.
Each bar comes from a different chip generation, model shape, and sequence range. They should not be averaged together into one "STEEL is Nx faster" claim.
Fusion avoids writing intermediate A / P tensors to DRAM.
DATO's compile time blows up past 4096 tokens — the 9.6× SotA claim lives entirely in the shortest, easiest regime STEEL tests elsewhere.
The paper's agentic-OS framing implies broad coverage. The evaluation matrix is narrower than the framing.