Interactive Visualization Episode Companion arXiv: 2512.12949 Hopper DSM FFN Fusion

FlashFuser and Hopper-Era FFN Kernel Fusion

The visual story here is not “fusion is good.” It is that Hopper’s cluster-scoped shared memory and inter-SM communication change the size of the fusible region. These diagrams track where activations live, when DSM helps, and why large FFN intermediates stop being forced into an HBM round-trip so early.

V100-era A100 H100 Hopper+ FP16 compute rising faster HBM bandwidth rises slowly ~1000 TFLOPS ~3 TB/s
FFN Time Share
40–60%
Dense transformer feed-forward blocks can dominate inference slices at seq=512.
Kernel Claim
3.3×
Mocked from paper-reported scale: kernel wins are much larger than end-to-end wins.
On-Chip Goal
DSM > HBM
Spill later, across a larger on-chip pool shared by a thread-block cluster.

Single-SM fusion hits a wall; cluster fusion moves the wall

Left: classic FFN operator chaining spills to HBM when the hidden activation no longer fits one SM’s scratchpad. Right: FlashFuser treats a Hopper cluster as a larger schedulable workspace.

Compute stage On-chip exchange / DSM HBM spill

References and Episode Links

FlashFuser

Ziyu Huang et al., 2025. Primary paper behind this page’s DSM-aware FFN fusion story.

arXiv: 2512.12949

TVM

Tianqi Chen et al., 2018. Compiler-first framing for graph and schedule optimization.

Scholar link

FlashAttention

Tri Dao et al., 2022. IO-awareness changed how people reason about bytes moved versus FLOPs spent.

Scholar link

Hopper Benchmarking

Weile Luo et al., 2024. Useful backdrop for thread-block clusters, DSM, and Hopper-era memory behavior.

Scholar link

Chimera

Compute-intensive fusion before the cluster-level leap. Good contrast point for what FlashFuser adds.

Scholar link

Related Episodes

Prior AI Post Transformers episodes on serving math, speculative decoding, phase-split inference, and long-sequence efficiency.

Browse podcast site

Additional arXiv IDs explicitly present in the transcript text: none beyond 2512.12949.