Interactive Visualization · FPGA trigger ML companion

Fast FPGA BDT Inference for LHC Triggers

A visual tour of what happens when a 100-tree, depth-4 boosted decision tree stops being software and becomes an ultra-low-latency decision circuit on a Xilinx VU9P.

arXiv:2002.02534 5-class jet tagging 100 trees · depth 4 200 MHz · II = 1 Sub-100 ns
Target fabricXilinx VU9P FPGA
Event cadence25 ns bunch crossings
Feature vector16 jet features
Precision window11–15 bit fixed-point
Comparison anchor3-layer MLP baseline
Transcript arXiv IDs detected: 2002.02534

Latency Budget as Geometry

The trigger problem is about arranging nanoseconds and hardware blocks, not maximizing software throughput. This view shows where the BDT sits in the path from collision to accept/reject logic.

Tab 1 / 4
streamed signal path hardwired BDT fabric latency pressure

Why trees survive here

Threshold comparisons and branch routing map naturally to LUT fabric. The BDT spends logic on control and accumulation, while a dense MLP spends more of its budget on multiply-accumulate units and DSP blocks.

Clock
200 MHz
Period
5 ns
Pipeline
II = 1
Model block
~60 ns

Hover any stage to inspect local timing and resource emphasis.

From Tree to Circuit

Each split becomes a comparator, each branch becomes signal routing, and each leaf becomes a small fixed-point score contribution. Step through the conversion and inspect the path logic.

Tab 2 / 4

Hardwired inference

No pointer chasing, no software walk. The tree is unrolled so traversal becomes combinational signal flow, then class scores are summed across the ensemble.

Depth
4
Leaves/tree
16
Classes
5
Estimators
100

Hover comparators and leaves to reveal the fixed-point threshold or score value.

Quantization Window

The paper’s useful operating region is not “full precision or bust.” This section shows a mock but paper-shaped view of class retention as bit width shifts from fragile to comfortable.

Tab 3 / 4

Mid-teens bits are the calm zone

At very low precision, thresholds and leaf outputs drift enough to bend operating points. In the 11–15 bit region, the model still behaves much like the floating-point reference while staying inside a tolerable hardware budget.

Best AUC
95.3%
Weakest class
Quark
Stable zone
11–15b
Failure edge
≤ 8b

Switch the overlay to view retention, threshold jitter, or resource savings.

Accuracy vs Fabric Pressure

The BDT does not win by raw classifier quality. It wins by landing in a different resource shape: more LUT-heavy, less DSP-hungry, and still fast enough to fit the trigger window.

Tab 4 / 4

Different hardware dialects

The BDT burns more of the configurable logic fabric on decisions and reductions. The MLP concentrates more demand in multipliers and DSP blocks. Which is better depends on the board budget around the model, not on a single headline metric.

BDT AUC avg
92.9%
MLP edge
+0.9 pts
BDT DSP
Low
BDT LUT
Higher

Toggle between model metrics and a frequency-depth pressure map.

References

Selected sources behind the visuals and episode framing.

1
Fast inference of Boosted Decision Trees in FPGAs for particle physics
Summers et al., 2020
arXiv
2
Greedy Function Approximation: A Gradient Boosting Machine
Friedman, 2001
Scholar
3
Fast inference of deep neural networks in FPGAs for particle physics
Duarte et al., 2018
Scholar
4
Efficient, reliable and fast high-level triggering using a bonsai boosted decision tree
Gligorov, Williams, 2013
Scholar
5
Boosted Decision Trees in the Level-1 Muon Endcap Trigger at CMS
CMS Collaboration, 2018
Scholar
6
Low latency transformer inference on FPGAs for physics applications with hls4ml
Jiang et al., 2025
Scholar