This page is about timing pressure, not prose. The paper’s core move is to compress a jet classifier into a deterministic FPGA block that can sit inside a Level-1 trigger path: batch-one, fixed-point, hardware-aware, and small enough that routing, DSP pressure, and memory matter as much as classifier accuracy.
Step through the actual placement of the neural block. The model is only one island in a larger hard-real-time path, which is why the paper’s strongest claim is about a compact classifier block rather than the whole detector pipeline.
The FPGA does not “run a model” the way a GPU runs software. It instantiates a schedule of multipliers, adders, registers, and memory accesses. Use the controls to move between aggressive parallelism and hardware reuse.
These charts use realistic mock numbers derived from the paper’s story: shrinking precision and changing reuse move latency and resource footprint far more sharply than they move classification quality, which is exactly why hardware-aware model selection matters.
In this workload, “faster” and “bigger” are almost the same decision. Wider layers and lower reuse burn parallel hardware to cut cycles; fixed-point arithmetic claws some of that budget back.
This heatmap treats FPGA deployment as a search problem over hardware-aware model variants. Hover any cell: the best-looking software model is not always the one that survives the routing and latency envelope.
hls4ml matters because it turns these cells from hand-wired firmware projects into a tractable iteration loop: compile, measure, prune, quantize, retrain, and re-synthesize.