DFX: Multi-FPGA Acceleration for Transformer Inference
A visual companion to the episode on why low-batch autoregressive decode can leave GPUs underfilled, how DFX maps GPT-style generation across four FPGAs, and where the paper’s large latency and efficiency claims depend on baseline choices.
Paper-reported speedup over the four-V100 baseline.
3.99×
Paper-reported energy efficiency gain in the same setup.
8.21×
Paper-reported cost effectiveness under the authors’ comparison.
Prefill is wide. Decode is serial.
The paper calls them summarization and generation. This view uses modern prefill/decode language and shows why single-token serving creates a utilization cliff.
Stage Occupancy Heatmap
Hover cells to inspect token-parallel pressure across layers and time
The message is structural, not benchmark-specific: prefill amortizes large matrix work across many prompt tokens, decode repeats dependent single-token steps.
Narrative Translation
2022 wording mapped to 2026 serving language
The paper’s “summarization stage” aligns with prefill. Its “generation stage” aligns with decode. That historical rename matters because modern serving papers optimize those phases separately.
One token forces the next
This stepper walks the autoregressive loop. The workload shrinks from many prompt tokens to one new token, while the cache and layer stack persist.
Autoregressive Serving Timeline
Step-by-step SVG diagram with animated emphasis
Decode is not one big GEMM. It is a stream of dependent passes through attention, MLP, and projection, with the KV cache carrying context forward.
KV Cache Pressure
Illustrative matrix access pattern
Older papers often under-name the cache. The bottleneck still exists: decode keeps touching stored key/value state while adding tiny increments of new compute.
Per-Token Critical Path
Mock cycle budget for one generated token
The point is latency shape, not exact counts: every next token waits for the previous one’s logits, sample, and cache update.
Four FPGAs as a coordinated decode appliance
DFX is not just “attention on FPGA.” It claims end-to-end execution across embedding, projections, attention, feed-forward, normalization, residuals, and LM head, with HBM-aware tiling and model-parallel placement.
Appliance Topology
Switch between logical pipeline and HBM traffic emphasis
compute stagecross-device handoffHBM-heavy region
Layer Residency Map
Illustrative hardware-aware sharding
Model parallelism itself is not novel. The novelty claim is the FPGA-specific execution schedule and memory placement for low-latency generation.
End-to-End Coverage
What the paper says stays on the appliance
The systems distinction from many earlier FPGA papers is that DFX argues for fewer host escapes and fewer isolated “accelerated subgraphs.”
Headline gains, plus the fairness question
The paper reports large wins against four V100s. This tab keeps those numbers visible, then lets you compare them against alternate serving assumptions rather than treating one GPU baseline as timeless.
Relative Advantage Dashboard
Paper baseline versus hypothetical stronger GPU serving stacks
Only the first mode reproduces the paper’s reported deltas. The others are mock sensitivity views that visualize the episode’s skepticism: stronger scheduling and batching can compress the margin.
Regime Map
Where specialization looks strongest
The episode’s durable claim is regime-specific: low-batch, latency-sensitive decode is a different target from throughput-oriented prompt processing.
Context Timeline
Papers that moved the serving conversation later
Speculative decoding, KV compression, GQA, and prefill/decode disaggregation change what an “apples-to-apples” comparison looks like after 2022.