Skip to main navigation Skip to search Skip to main content

Highly Parallel CNN Accelerator for RepVGG-Like Network Training on FPGAs

  • Chuliang Guo
  • , Binglei Lou
  • , David Boland
  • , Philip H.W. Leong*
  • *Corresponding author for this work
  • Zhejiang University
  • School of Electronics and Information Engineering, Harbin Institute of Technology
  • The University of Sydney

Research output: Contribution to journalArticlepeer-review

Abstract

In this article, we propose a generic FPGA-based training accelerator tailored for RepVGG-like networks, which strikes a balance between maximizing training-time accuracy and minimizing inference-time latency. The proposed accelerator leverages fine-grain channel-level parallelism within computational units specially designed for multiple branches of the basic building block within the RepVGG-like network. Specifically, we employ a Conv block for forward Conv and backward deConv, along with a dilated Conv block, including a weight kernel partition scheme for efficient weight gradient calculation. Furthermore, we aggressively exploit a 2-stage coarse-grain task-level parallelism for low-latency CNN training: 1) parallelism among multiple branches of the basic building block of RepVGG and 2) parallelism between error back-propagation and weight gradient calculation in the backward path. Through experiments on the CIFAR-10 dataset using 16-bit fixed-point arithmetic, we demonstrate state-of-the-art batch 1 throughput of 150 GOPs for training and 183 GOPs for inference.

Original languageEnglish
Pages (from-to)554-558
Number of pages5
JournalIEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems
Volume44
Issue number2
DOIs
StatePublished - 2025
Externally publishedYes

Keywords

  • Convolutional neural network (CNN) training
  • FPGA
  • HLS
  • RepVGG

Fingerprint

Dive into the research topics of 'Highly Parallel CNN Accelerator for RepVGG-Like Network Training on FPGAs'. Together they form a unique fingerprint.

Cite this