Return
A Groupwise Add--Multiply--Shift--Accumulate Datapath for Efficient DNN Accelerators
DOI:10.1109/TVLSI.2026.3689624.png)
Abstract
En 中文
Multiply–accumulate(MAC) units account for a large fraction of the power and area in modern deep neural network (DNN) accelerators. Although low-bitwidth quantization reduces hardware overhead, the high cost of multipliers remains a fundamental bottleneck in modern accelerator datapaths. This article proposes add–multiply–shift–accumulate (AMC), a groupwise arithmetic datapath that reduces multiplier count by sharing base multiplications across groups of neighboring weights and generating residual products using lightweight shift operation. To support efficient deployment, we design a compact residual encoding and buffer organization that allows AMC arrays to be constructed with minimal decoding and control overheads. While AMC can be directly applied to existing quantized models, we further introduce a lightweight residual-aware fine-tuning (RAF) procedure to increase AMC compatibility. We implement AMC-based accelerators in SystemVerilog and synthesize them in TSMC 28-nm CMOS technology across operating frequencies from 500 MHz to 1 GHz. At the compute unit level, AMC reduces arithmetic area by 39.5%–62.8% and dynamic power by 32.2%–60.3% compared with optimized baseline multipliers. When integrated into CNN and Vision Transformer accelerators, AMC achieves $1.34\times $ – $18.90\times $ higher area efficiency and up to $10.16\times $ higher energy efficiency than prior designs while preserving baseline inference accuracy.
Keywords:
Deep learning
hardware–software codesign
multiply–accumulate
weight distribution
Journal
I
IF:
3.1
Papers:
440
Citations:
7.3K

