A whole-network FPGA story from the AlexNet-to-VGG era: unify convolution and fully connected layers behind one matrix engine, then win or lose on memory traffic, reuse, and burst scheduling rather than raw multiply counts alone.
Compact trail of papers and adjacent episodes behind the visuals.