AI Post Transformers · Episode Companion

AMD XDNA NPU Sparsity-Aware Attention for Long-Context Prefill

STEEL fuses FlashAttention-2's tiling into a three-stage pipeline purpose-built for AMD's spatial-dataflow XDNA NPU, and fixes the load imbalance the causal mask creates on a fixed pipeline by scattering rows across pipelines instead of assigning contiguous chunks.

arXiv:2607.09385 Victor J. B. Jung et al. · 2026 AMD Research · ETH Zürich · U. Bologna Interactive Viz ↗

Two Execution Models

A GPU is SIMT with a cache hierarchy deciding what stays close to compute. XDNA is a 2-D grid of VLIW tiles wired by an on-chip network — the compiler, not a cache, decides which tile does what and schedules every transfer.

The Three-Stage Fused Pipeline

STEEL splits FlashAttention-2's math across three dedicated AIE cores connected by IRON ObjectFIFOs — typed, synchronized queues — instead of one core looping through every step serially.

3
dedicated AIE cores per pipeline
10
parallel pipelines on-chip
0
DRAM round-trips for A / P tensors

Causal Mask Load Imbalance

Roughly half the attention matrix is masked to zero. Assigning contiguous row-blocks per pipeline gives early pipelines mostly-masked (cheap) tiles and late pipelines full unmasked (expensive) tiles — since K/V broadcast waits for every consumer, the slowest pipeline stalls the whole array.

Mem-Tile Port Budget

Each Mem tile has only 6 ports. Every STEEL pipeline needs 4: one distributed Q tile, one collected O tile, and two for swizzling intermediate P tiles between the softmax and PV stages. K/V are broadcast, shared across all pipelines.

42 / 48
Mem-tile ports consumed across 10 pipelines
4
ports per pipeline (Q, O, 2× P-swizzle)
6
ports headroom left on the whole chip

Headline Numbers — Three Different Benchmarks

Each bar comes from a different chip generation, model shape, and sequence range. They should not be averaged together into one "STEEL is Nx faster" claim.

Off-Chip Data Movement @ seq len 4096

Fusion avoids writing intermediate A / P tensors to DRAM.

DATO Comparison vs STEEL's Own Range

DATO's compile time blows up past 4096 tokens — the 9.6× SotA claim lives entirely in the shortest, easiest regime STEEL tests elsewhere.

What's Actually Tested

The paper's agentic-OS framing implies broad coverage. The evaluation matrix is narrower than the framing.

Tested / covered Not tested / open gap
Every benchmark in the paper — layer-by-layer, DATO, and the energy sweep — is prefill-only, MHA-only, causal-mask-only, single-vendor. Decode, GQA, and paged-KV-cache integration are not evaluated anywhere in this work.

References