AI Post Transformers / Visual Companion

Caffeine: A Unified FPGA for CNNs

A whole-network FPGA story from the AlexNet-to-VGG era: unify convolution and fully connected layers behind one matrix engine, then win or lose on memory traffic, reuse, and burst scheduling rather than raw multiply counts alone.

Paper context
Published
2016
Key claim
1 FPGA / whole CNN
Main tension
Compute vs bandwidth
This page uses schematic, normalized values to visualize the paper’s argument: once convolution improves, fully connected layers can become the dominant data-movement bottleneck.

References

Compact trail of papers and adjacent episodes behind the visuals.