FP-DNN (2017)
Hybrid RTL-HLS templates for mapping TensorFlow-era DNNs to FPGA inference hardware.
paper PDF · DOI
A visual tour of FP-DNN’s central wager: squeeze convolutions, fully connected layers, and parts of LSTM execution into one reusable matrix core, then wrap it with just enough programmable glue to keep deployment from becoming a handcrafted hardware exercise.
The pressure curve was simple: deeper CNNs and recurrent models kept growing, but the hardware flow for FPGAs still looked manual. FP-DNN tries to move effort from per-model hardware design into a compiler and template stack.
The page’s central picture is the flow from TensorFlow graph parsing to scheduling, template selection, bitstream generation, and runtime control. Click the step buttons to reveal where FP-DNN leans on hand-tuned RTL versus HLS-generated orchestration.
FP-DNN’s big abstraction move is to normalize different layers into matrix multiplication as often as possible. Switch workloads to see how the dataflow pattern changes while the central compute shape stays familiar.
The paper’s pitch combines throughput, latency, and energy efficiency. Toggle the comparison lens below: the chart changes, but the fairness warning remains because platform, batching, and precision choices are doing real work in these outcomes.
This view places FP-DNN in a longer hardware-compiler arc, from earlier template systems to modern dynamic-shape compilers and transformer-oriented mixed-precision accelerators. The scatter plot is paired with an architecture slice showing where communication efficiency dominates.
Hybrid RTL-HLS templates for mapping TensorFlow-era DNNs to FPGA inference hardware.
paper PDF · DOI
Automatic generation of FPGA learning accelerators for neural network families.
scholar link
Automation push toward high-performance DNN hardware generation on FPGAs.
scholar link
Uniform representation and acceleration strategy for CNNs on FPGAs.
scholar link
Compiler-era reminder that dynamic shapes and whole-graph optimization become first-class later.
scholar link
Modern FPGA acceleration shifts toward transformable arithmetic, memory movement, and finer precision control.
TATAA · 2025 mixed precision