arrow
Return

Partial Sum Quantization for Computing-In-Memory-Based Neural Network Accelerator

delete2023-08-01
delete5
PRE
AI
J
Jinyu Bai
W
Wenlu Xue
Y
Yunqian Fan
S
Sifan Sun
W
Wang Kang *
DOI:10.1109/TCSII.2023.3246562delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Computing-in-memory (CIM) has been successful as an ideal hardware platform to improve the performance and efficiency of convolutional neural networks (CNNs). However, owing to the limited size of a memory array, the input and weight matrices of a convolution operation have to be split into sub-matrices, involving partial sums. Generally, high-resolution analog-to-digital converters (ADCs) are used to obtain partial sums for maintaining the computing precision, but at the cost of high area and energy. Partial sum quantization (PSQ), which can be exploited to significantly reduce the ADC's resolution, is still an open question in this field. This brief proposes a novel PSQ approach for CIM using post-training quantization based on a newly defined array-wise granularity. Meanwhile, as the non-linearity of ADCs' transfer function has a severe impact on the accuracy, a gradient estimation method based on smooth approximation is proposed to solve such a problem. Experiments on various CNNs show that the required ADCs' resolution can be reduced from 11-bit to even 3-bit with slight accuracy loss (similar to 1.63%), and the energy-efficiency is increased by up to 224%.
Keywords:
Computing-in-memory
partial sum quantiza-tion
post-training quantization
gradient estimation

Journal

I
IEEE Transactions on Circuits and Systems and Express Briefs
IF:
4.9
Papers:
8.8K
Citations:
2.5W

Organization

B
Beihang University
Scholars:
5.1W
Papers: 4.1W
Citations: 37