AI Post Transformers · Interactive Visualization

Automating CNN Mapping on Embedded FPGAs

fpgaConvNet reframes CNN deployment as a schedulable streaming graph: layers become actors, activations become tokens, and hardware choices become search variables. This page turns that claim into visual machinery: dataflow graphs, schedule heatmaps, resource frontiers, and bottleneck maps.

arXiv:1711.08740 Venieris & Bouganis · 2017 CNN → SDF → FPGA Toolflow Known arXiv IDs: 1711.08740 Extra IDs from transcript: none detected
CNN LAYERS → SDF GRAPH → FPGA formal schedule + buffer sizing + resource tradeoffs

Toolflow as a Visual Control Surface

Instead of tuning one handcrafted accelerator, fpgaConvNet treats the architecture as a parameterized graph. This view shows the journey from model topology to partitioning, scheduling, and generated hardware.

Why This Matters

Embedded inference is constrained by movement of activations and weights, not just multiply-accumulates. Formal dataflow exposes where buffering, reconfiguration, and parallelism help or hurt.

What Changes Per Mode

Throughput mode tolerates batching and phase changes to keep the pipe full. Latency mode prefers fewer disruptive reconfigurations and shorter per-input paths.

Interactive Readout

Layer graph view: the CNN is still recognizable, but each layer has already become a hardware actor with token rates and implementation choices attached.

Network stages
8
Candidate partitions
3
Analytic buffers
14
Objective knobs
4

Synchronous Dataflow, Drawn as Scheduling Physics

Actors fire at fixed token rates. That constraint is severe, but it enables static schedules and buffer sizing. Hover any cell to inspect occupancy, utilization, or burst pressure.

compute actor merge / branch actor high occupancy reload or bandwidth stress

Design-Space Exploration Instead of One Benchmark Story

The paper’s strongest claim is not a single speedup number. It is that the optimizer can move across CNN topologies and objectives. These charts use mock but plausible values to show how the frontier bends when the goal changes.

Fair Comparison Caveat

“Same power budget” does not guarantee equal software tuning, precision, or batching assumptions. A headline speedup is only as clean as the deployment match.

Generality Claim

Static CNNs fit SDF well. Residual, dense, and branchy models can still be mapped, but often through partition choices and compile-time shaping rather than full abstraction magic.

What The Solver Sees

Parallelism factors, weight reload strategy, partitioning count, and on-chip memory use become numeric decisions. The hardware becomes searchable.

Where the Elegant Model Meets Ugly Hardware

This section visualizes the skepticism. The optimizer may prefer a point that looks good analytically, but routing congestion, BRAM fragmentation, and off-chip bandwidth can bend reality away from the neat frontier.

Risk Map

Baseline concern: the farther the design leans into memory reloads and dense connectivity, the more the analytic model risks underestimating practical cost.

2017 vs Later Work

Later FPGA systems for transformers, early-exit networks, and compute-in-memory make the schedule less static and the memory model even more central. The old SDF advantage becomes conditional.

Visual Heuristic

Green regions are schedule-friendly. Orange and red regions indicate design points where a compile-time graph is still useful, but no longer sufficient as a complete predictor.

Place/route risk
0.42
Bandwidth strain
0.56
Model drift
0.31
Validation coverage
0.22

References

fpgaConvNet: A Toolflow for Mapping Diverse Convolutional Neural Networks on Embedded FPGAsVenieris, Bouganis, 2017 · arXiv:1711.08740
Static Scheduling of Synchronous Data Flow Programs for DSPLee, Messerschmitt, 1987 · Scholar
Synchronous Data FlowLee, Messerschmitt, 1987 · Scholar
Scenario-aware DataflowStuijk, Geilen, Theelen, Basten, 2011 · Scholar
DNNWeaverSharma et al., 2016 · Scholar
CaffeineZhang et al., 2016 · Scholar
FINNUmuroglu et al., 2017 · Scholar
Efficient Processing of Deep Neural NetworksSze et al., 2017 · Scholar
ViTA / ME-ViT / DRViTEdge FPGA accelerators for transformer inference, 2023–2025 · ViTA · ME-ViT · DRViT
Dynamic and Memory-Centric Follow-onsEarly-exit FPGA hardware and compute-in-memory, 2024–2025 · Early Exit · CIM Review
Related EpisodeAI Post Transformers: FPGA Neural Network Accelerators for Space · Listen