FCCM 2017 TensorFlow graph → FPGA bitstream RTL + HLS hybrid templates

Automating DNN Compilation for FPGA Accelerators

A visual tour of FP-DNN’s central wager: squeeze convolutions, fully connected layers, and parts of LSTM execution into one reusable matrix core, then wrap it with just enough programmable glue to keep deployment from becoming a handcrafted hardware exercise.

Why This Paper Was Timely

The pressure curve was simple: deeper CNNs and recurrent models kept growing, but the hardware flow for FPGAs still looked manual. FP-DNN tries to move effort from per-model hardware design into a compiler and template stack.

Compiler Pipeline and Reuse Strategy

The page’s central picture is the flow from TensorFlow graph parsing to scheduling, template selection, bitstream generation, and runtime control. Click the step buttons to reveal where FP-DNN leans on hand-tuned RTL versus HLS-generated orchestration.

Operator-to-Core Mapping and Tile Heatmaps

FP-DNN’s big abstraction move is to normalize different layers into matrix multiplication as often as possible. Switch workloads to see how the dataflow pattern changes while the central compute shape stays familiar.

Performance Story, Precision Story, and Fairness Story

The paper’s pitch combines throughput, latency, and energy efficiency. Toggle the comparison lens below: the chart changes, but the fairness warning remains because platform, batching, and precision choices are doing real work in these outcomes.

Programmability vs Specialization Across Generations

This view places FP-DNN in a longer hardware-compiler arc, from earlier template systems to modern dynamic-shape compilers and transformer-oriented mixed-precision accelerators. The scatter plot is paired with an architecture slice showing where communication efficiency dominates.

Reference Trail

FP-DNN (2017)

Hybrid RTL-HLS templates for mapping TensorFlow-era DNNs to FPGA inference hardware.
paper PDF · DOI

DeepBurning (2016)

Automatic generation of FPGA learning accelerators for neural network families.
scholar link

DNNBuilder (2018)

Automation push toward high-performance DNN hardware generation on FPGAs.
scholar link

Caffeine (2016)

Uniform representation and acceleration strategy for CNNs on FPGAs.
scholar link

BladeDISC (2023)

Compiler-era reminder that dynamic shapes and whole-graph optimization become first-class later.
scholar link

TATAA and Mixed Precision FPGA Work

Modern FPGA acceleration shifts toward transformable arithmetic, memory movement, and finer precision control.
TATAA · 2025 mixed precision