AI Post Transformers • Interactive Visualization

DFX: Multi-FPGA Acceleration for Transformer Inference

A visual companion to the episode on why low-batch autoregressive decode can leave GPUs underfilled, how DFX maps GPT-style generation across four FPGAs, and where the paper’s large latency and efficiency claims depend on baseline choices.

arXiv:2209.10797 2022 paper 4× Xilinx U280 vs 4× NVIDIA V100
5.58×
Paper-reported speedup over the four-V100 baseline.
3.99×
Paper-reported energy efficiency gain in the same setup.
8.21×
Paper-reported cost effectiveness under the authors’ comparison.

Prefill is wide. Decode is serial.

The paper calls them summarization and generation. This view uses modern prefill/decode language and shows why single-token serving creates a utilization cliff.

Stage Occupancy Heatmap
Hover cells to inspect token-parallel pressure across layers and time
idle / narrow work busy / wider parallel work near saturation
Normalized Phase Mix
Mock workloads for latency-oriented serving
The message is structural, not benchmark-specific: prefill amortizes large matrix work across many prompt tokens, decode repeats dependent single-token steps.
Narrative Translation
2022 wording mapped to 2026 serving language
The paper’s “summarization stage” aligns with prefill. Its “generation stage” aligns with decode. That historical rename matters because modern serving papers optimize those phases separately.

One token forces the next

This stepper walks the autoregressive loop. The workload shrinks from many prompt tokens to one new token, while the cache and layer stack persist.

Autoregressive Serving Timeline
Step-by-step SVG diagram with animated emphasis
Decode is not one big GEMM. It is a stream of dependent passes through attention, MLP, and projection, with the KV cache carrying context forward.
KV Cache Pressure
Illustrative matrix access pattern
Older papers often under-name the cache. The bottleneck still exists: decode keeps touching stored key/value state while adding tiny increments of new compute.
Per-Token Critical Path
Mock cycle budget for one generated token
The point is latency shape, not exact counts: every next token waits for the previous one’s logits, sample, and cache update.

Four FPGAs as a coordinated decode appliance

DFX is not just “attention on FPGA.” It claims end-to-end execution across embedding, projections, attention, feed-forward, normalization, residuals, and LM head, with HBM-aware tiling and model-parallel placement.

Appliance Topology
Switch between logical pipeline and HBM traffic emphasis
compute stage cross-device handoff HBM-heavy region
Layer Residency Map
Illustrative hardware-aware sharding
Model parallelism itself is not novel. The novelty claim is the FPGA-specific execution schedule and memory placement for low-latency generation.
End-to-End Coverage
What the paper says stays on the appliance
The systems distinction from many earlier FPGA papers is that DFX argues for fewer host escapes and fewer isolated “accelerated subgraphs.”

Headline gains, plus the fairness question

The paper reports large wins against four V100s. This tab keeps those numbers visible, then lets you compare them against alternate serving assumptions rather than treating one GPU baseline as timeless.

Relative Advantage Dashboard
Paper baseline versus hypothetical stronger GPU serving stacks
Only the first mode reproduces the paper’s reported deltas. The others are mock sensitivity views that visualize the episode’s skepticism: stronger scheduling and batching can compress the margin.
Regime Map
Where specialization looks strongest
The episode’s durable claim is regime-specific: low-batch, latency-sensitive decode is a different target from throughput-oriented prompt processing.
Context Timeline
Papers that moved the serving conversation later
Speculative decoding, KV compression, GQA, and prefill/decode disaggregation change what an “apples-to-apples” comparison looks like after 2022.
References

Cited papers and related episode context

DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation Hong et al., 2022. Primary source for the four-FPGA appliance, GPT-style decode focus, and the reported 5.58× / 3.99× / 8.21× deltas.
GPipe Training-era pipeline parallelism context. Useful for seeing that sharding itself predates DFX.
Megatron-LM Model-parallel lineage for large transformers. DFX differs in hardware target and decode-first scheduling.
FTRANS Earlier FPGA acceleration context. The episode contrasts subcomponent acceleration with DFX’s end-to-end claim.
Speculative Decoding Later algorithmic shift that can reduce serial decode pain without changing the base hardware.
Prefill/Decode Aggregation or Disaggregation? Later systems work that formalizes the phase split DFX observed early.
AI Post Transformers: Speculative Decoding in Real vLLM Serving Episode context for why serving-stack quality can matter as much as raw accelerator specialization.
AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design Later decode-oriented framing that makes DFX feel historically early rather than obsolete.