Interactive Visualization AI Post Transformers Detector ML

Fast FPGA Inference for LHC Triggers

This page is about timing pressure, not prose. The paper’s core move is to compress a jet classifier into a deterministic FPGA block that can sit inside a Level-1 trigger path: batch-one, fixed-point, hardware-aware, and small enough that routing, DSP pressure, and memory matter as much as classifier accuracy.

arXiv 1804.06913 Input 16 engineered jet features Classes quark, gluon, W, Z, top Latency ~75 ns classifier block Transcript IDs 1804.06913 only
Collision Rate
40 MHz
Events arrive on a clock, not on analyst time.
Trigger Reality
< 1 μs
Worst-case latency and jitter dominate design.
Deployment Style
Batch-1
No large-batch amortization. Every event stands alone.
Workflow
ML → HLS → FPGA
hls4ml turns trained networks into synthesis-ready hardware.

From Proton Collision to Trigger Bit

Step through the actual placement of the neural block. The model is only one island in a larger hard-real-time path, which is why the paper’s strongest claim is about a compact classifier block rather than the whole detector pipeline.

Selected Stage

Latency Budget Split

Legend

detector / stream feature engineering classifier accept / reject logic

Dense Network Becomes a Circuit Template

The FPGA does not “run a model” the way a GPU runs software. It instantiates a schedule of multipliers, adders, registers, and memory accesses. Use the controls to move between aggressive parallelism and hardware reuse.

Current Mapping

Why It Changes

Feature Vector

Accuracy Is Not the Only Axis

These charts use realistic mock numbers derived from the paper’s story: shrinking precision and changing reuse move latency and resource footprint far more sharply than they move classification quality, which is exactly why hardware-aware model selection matters.

Selected Regime

Key Signal

Classifier latency75 ns
Resource pressuremedium
Relative quality0.93 AUC

Interpretation

In this workload, “faster” and “bigger” are almost the same decision. Wider layers and lower reuse burn parallel hardware to cut cycles; fixed-point arithmetic claws some of that budget back.

Design Surface: Width × Precision × Reuse

This heatmap treats FPGA deployment as a search problem over hardware-aware model variants. Hover any cell: the best-looking software model is not always the one that survives the routing and latency envelope.

Hovered Design

What This Section Says

hls4ml matters because it turns these cells from hand-wired firmware projects into a tractable iteration loop: compile, measure, prune, quantize, retrain, and re-synthesize.

Color Scale

cool / easy tense hard limit

References

1
Fast inference of deep neural networks in FPGAs for particle physics — Duarte et al., 2018
2
A Survey on Performance Optimization of High-Level Synthesis Tools — Huang et al., 2020
3
FINN-R: An End-to-End Deep-Learning Framework for Fast Exploration of Quantized Neural Networks — Blott et al., 2018
4
Fast convolutional neural networks on FPGAs with hls4ml — Aarrestad et al., 2021
5
Project Brainwave: Serving DNNs in Real Time at Datacenter Scale — Chung et al., 2018
6
AI Post Transformers: FPGA Neural Network Accelerators for Space — Hal Turing & Dr. Ada Shannon, 2026