arrow
Return

CDCA: a column-serial and dual-engine collaborative accelerator for depthwise separable convolution

delete2026-08-09
delete0
PRE
AI
Y
Yiming Ouyang
Y
Yuanzhen Jia
J
Jianhua Li
Q
Qi Wang *
DOI:10.1007/s11227-026-08752-7delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Depthwise separable convolution (DSC) effectively reduces computational complexity in convolutional neural networks. However, the heterogeneous characteristics of its constituent operators often lead to low hardware utilization and pipeline stalls in traditional accelerators. To address this, we propose a column-serial and dual-engine collaborative accelerator (CDCA) on FPGA. This design aims to improve energy efficiency through a dataflow-driven architecture. Firstly, we introduce a decoupled heterogeneous dual-engine architecture. This design tailors optimizations to the distinct patterns of different operators, enabling efficient collaboration and significantly higher hardware utilization. Secondly, we propose a column-serial sliding-window dataflow to overcome the bottleneck of memory-bound operators. By using shift-register chains, the arriving data can be processed immediately, maximizing on-chip data reuse and reducing redundant memory accesses typical of traditional line-buffer architectures. Finally, we implement a tile-based partitioning strategy combined with double-buffering, which enables effective overlap between computation and data transfer. We evaluated the proposed accelerator on a Xilinx ZCU104 platform running at 100 MHz, using MobileNetV1 as the main benchmark. CDCA achieves an end-to-end inference throughput of 29.50 FPS and an achieved performance of 20.42 GOPS, with a measured FPGA power consumption of only 1.93 W during MobileNetV1 inference, corresponding to an energy efficiency of 10.58 GOPS/W. These results show that CDCA provides a low-power FPGA design point for DSC acceleration, making it suitable for resource-constrained edge inference scenarios.
Keywords:
Depthwise separable convolution
Dual-engine architecture
Column-serial sliding
Time-division multiplexing
FPGA accelerator

Journal

Journal of Supercomputing cover
Journal of Supercomputing
IF:
2.7
Papers:
990
Citations:
1.0W

Organization

S