fpgaConvNet reframes CNN deployment as a schedulable streaming graph: layers become actors, activations become tokens, and hardware choices become search variables. This page turns that claim into visual machinery: dataflow graphs, schedule heatmaps, resource frontiers, and bottleneck maps.
Instead of tuning one handcrafted accelerator, fpgaConvNet treats the architecture as a parameterized graph. This view shows the journey from model topology to partitioning, scheduling, and generated hardware.
Embedded inference is constrained by movement of activations and weights, not just multiply-accumulates. Formal dataflow exposes where buffering, reconfiguration, and parallelism help or hurt.
Throughput mode tolerates batching and phase changes to keep the pipe full. Latency mode prefers fewer disruptive reconfigurations and shorter per-input paths.
Layer graph view: the CNN is still recognizable, but each layer has already become a hardware actor with token rates and implementation choices attached.
Actors fire at fixed token rates. That constraint is severe, but it enables static schedules and buffer sizing. Hover any cell to inspect occupancy, utilization, or burst pressure.
The paper’s strongest claim is not a single speedup number. It is that the optimizer can move across CNN topologies and objectives. These charts use mock but plausible values to show how the frontier bends when the goal changes.
“Same power budget” does not guarantee equal software tuning, precision, or batching assumptions. A headline speedup is only as clean as the deployment match.
Static CNNs fit SDF well. Residual, dense, and branchy models can still be mapped, but often through partition choices and compile-time shaping rather than full abstraction magic.
Parallelism factors, weight reload strategy, partitioning count, and on-chip memory use become numeric decisions. The hardware becomes searchable.
This section visualizes the skepticism. The optimizer may prefer a point that looks good analytically, but routing congestion, BRAM fragmentation, and off-chip bandwidth can bend reality away from the neat frontier.
Baseline concern: the farther the design leans into memory reloads and dense connectivity, the more the analytic model risks underestimating practical cost.
Later FPGA systems for transformers, early-exit networks, and compute-in-memory make the schedule less static and the memory model even more central. The old SDF advantage becomes conditional.
Green regions are schedule-friendly. Orange and red regions indicate design points where a compile-time graph is still useful, but no longer sufficient as a complete predictor.