The visual story here is not “fusion is good.” It is that Hopper’s cluster-scoped shared memory and inter-SM communication change the size of the fusible region. These diagrams track where activations live, when DSM helps, and why large FFN intermediates stop being forced into an HBM round-trip so early.
FFN Time Share
40–60%
Dense transformer feed-forward blocks can dominate inference slices at seq=512.
Kernel Claim
3.3×
Mocked from paper-reported scale: kernel wins are much larger than end-to-end wins.
On-Chip Goal
DSM > HBM
Spill later, across a larger on-chip pool shared by a thread-block cluster.
Single-SM fusion hits a wall; cluster fusion moves the wall
Left: classic FFN operator chaining spills to HBM when the hidden activation no longer fits one SM’s scratchpad. Right: FlashFuser treats a Hopper cluster as a larger schedulable workspace.
Compute stageOn-chip exchange / DSMHBM spill
Reduce, shuffle, multiply, then spill only if forced
These are not abstract arrows. The active step highlights what the cluster communication pattern is doing to a large FFN intermediate tile.
Thread-block clusterDSM poolCurrent communication path
Where do activation tiles live as sequence length grows?
The grid shows a mock but plausible residency map for intermediate FFN tiles. Blue means safely on chip; hot orange-red means pressure forcing the tile toward HBM.
Kernel gains dilute on the way to service-level speedup
The paper’s shape of evidence is familiar: big kernel wins, smaller end-to-end wins. Toggle between speedup and traffic reduction to see how the bottleneck shifts across FFN variants.
FlashFuserTuned single-SM fused baselineService-level share of total latency
References and Episode Links
FlashFuser
Ziyu Huang et al., 2025. Primary paper behind this page’s DSM-aware FFN fusion story.