Abstract
In this article, we propose a generic FPGA-based training accelerator tailored for RepVGG-like networks, which strikes a balance between maximizing training-time accuracy and minimizing inference-time latency. The proposed accelerator leverages fine-grain channel-level parallelism within computational units specially designed for multiple branches of the basic building block within the RepVGG-like network. Specifically, we employ a Conv block for forward Conv and backward deConv, along with a dilated Conv block, including a weight kernel partition scheme for efficient weight gradient calculation. Furthermore, we aggressively exploit a 2-stage coarse-grain task-level parallelism for low-latency CNN training: 1) parallelism among multiple branches of the basic building block of RepVGG and 2) parallelism between error back-propagation and weight gradient calculation in the backward path. Through experiments on the CIFAR-10 dataset using 16-bit fixed-point arithmetic, we demonstrate state-of-the-art batch 1 throughput of 150 GOPs for training and 183 GOPs for inference.
| Original language | English |
|---|---|
| Pages (from-to) | 554-558 |
| Number of pages | 5 |
| Journal | IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems |
| Volume | 44 |
| Issue number | 2 |
| DOIs | |
| State | Published - 2025 |
| Externally published | Yes |
Keywords
- Convolutional neural network (CNN) training
- FPGA
- HLS
- RepVGG
Fingerprint
Dive into the research topics of 'Highly Parallel CNN Accelerator for RepVGG-Like Network Training on FPGAs'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver