Prefill and Decode Pull Different Bottlenecks
Toggle the regime. The visual shifts from prompt-wide attention work to repeated decode passes where the same MLP weights
are streamed for every new token.
Compute-dense region
HBM-heavy region
KV cache activity
Decode Pressure Profile
The paper’s framing is strongest when output length grows and effective batch stays small. Hover cells to inspect the workload shift.
What rises
Repeated weight traffic
What stays
Same dense SwiGLU path
Why fuse
Avoid dumping intermediates to HBM
SwiGLU Path: From Multi-Launch to Deep Fusion
Same math, different staging. The main intervention fuses up-projection, gate-projection, SiLU, and multiply into one decode-stage kernel.
Intermediate Residency Map
The heatmap tracks how much temporary activation material touches HBM at each step. Hotter means more off-chip traffic.
Cold / kept local
Warm / partial spill
Hot / heavy HBM
Decode Throughput Over Output Length
Mocked to match the reported shape: modest gains at short outputs, larger gains as decode length stretches and memory traffic dominates.
What the Stack-Level Result Suggests
The gain is not a headline algorithmic jump. It is the classic serving win: fewer intermediate reads and writes, better reuse, less launch fragmentation.
Best A100
+9.7%
Best H100
+13.2%
Interpretation
Strong for one decode regime
Profiler-Driven Kernel Selection
Instead of betting on one layout, the runtime probes candidates and chooses the winner for the target model shape, batch size, and GPU.
Why Row-Major and Column-Major Trade Places
Click a cell in the heatmap. The explainer redraws the favored reuse pattern, showing whether activation reuse or weight-tile reuse dominates.
Row-major tends to win
More token reuse, fatter batches
Column-major tends to win
Small-batch decode, giant weights
Paper lesson
Autotune, don’t hardcode lore