AI Post Transformers • viz companion

Deep Kernel Fusion for Transformer Decoding

A decode-centric systems story: when long generation turns token-by-token inference into a memory traffic problem, the SwiGLU MLP path stops being background machinery and becomes the roadblock.

arXiv:2602.11808
Posted Feb 12, 2026
Focus Decode-stage SwiGLU fusion
Stack SGLang + FlashInfer + CUDA Graphs
Hardware 4× A100 / 4× H100 SXM
Core Claim
MLP traffic matters in long decode
Reported Gain
Up to 9.7% A100 / 13.2% H100
Key Discipline
Fuse first stage, keep down-proj separate

Prefill and Decode Pull Different Bottlenecks

Toggle the regime. The visual shifts from prompt-wide attention work to repeated decode passes where the same MLP weights are streamed for every new token.
Compute-dense region HBM-heavy region KV cache activity

Decode Pressure Profile

The paper’s framing is strongest when output length grows and effective batch stays small. Hover cells to inspect the workload shift.
What rises
Repeated weight traffic
What stays
Same dense SwiGLU path
Why fuse
Avoid dumping intermediates to HBM

SwiGLU Path: From Multi-Launch to Deep Fusion

Same math, different staging. The main intervention fuses up-projection, gate-projection, SiLU, and multiply into one decode-stage kernel.

Intermediate Residency Map

The heatmap tracks how much temporary activation material touches HBM at each step. Hotter means more off-chip traffic.
Cold / kept local Warm / partial spill Hot / heavy HBM

Decode Throughput Over Output Length

Mocked to match the reported shape: modest gains at short outputs, larger gains as decode length stretches and memory traffic dominates.

What the Stack-Level Result Suggests

The gain is not a headline algorithmic jump. It is the classic serving win: fewer intermediate reads and writes, better reuse, less launch fragmentation.
Best A100
+9.7%
Best H100
+13.2%
Interpretation
Strong for one decode regime

Profiler-Driven Kernel Selection

Instead of betting on one layout, the runtime probes candidates and chooses the winner for the target model shape, batch size, and GPU.

Why Row-Major and Column-Major Trade Places

Click a cell in the heatmap. The explainer redraws the favored reuse pattern, showing whether activation reuse or weight-tile reuse dominates.
Row-major tends to win
More token reuse, fatter batches
Column-major tends to win
Small-batch decode, giant weights
Paper lesson
Autotune, don’t hardcode lore

References

arXiv IDs found in the provided material: 2602.11808. Related work below stays compact and link-first.
Deep Kernel Fusion for Transformer Decoding
Primary paper.
arXiv abstract · PDF
GLU Variants Improve Transformer
SwiGLU lineage and gated-FFN context.
Scholar
FlashAttention-2
Attention-side parallelism backdrop for why the MLP bottleneck becomes more visible.
Scholar
SGLang and PagedAttention
Serving runtime and KV-management context around the evaluated stack.
SGLang · PagedAttention
Welder and FlashInfer
Memory scheduling and inference-kernel engineering neighbors.
Welder · FlashInfer