1
Return

PQ-PIM: A pruning-quantization joint optimization framework for ReRAM-based-in DNN accelerator

delete2022-06-01
delete2
PRE
AI
Y
Yuhao Zhang
X
Xinyu Wang
X
Xikun Jiang
Y
Yuhan Yang
Z
Zhaoyan Shen
Z
Zhiping Jia *
DOI:10.1016/j.sysarc.2022.102531delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Pruning and quantization are two efficient techniques to achieve performance improvement and energy saving for ReRAM-based DNN accelerators. However, most existing ReRAM-based DNN accelerators using pruning and quantization are based on an overidealized multi-bit ReRAM crossbar while neglecting the practical structure constraints. Due to the restriction of immature process technology, the actual matrix-vector multiplication must be conducted in a smaller operation unit (OU) granularity with single bit ReRAM cells. In this paper, we propose an efficient pruning-quantization joint exploration framework for practical ReRAM-based DNN accelerator, termed as PQ-PIM, which consists of a patch-wise pruning-quantization algorithm based on patch importance analysis to compress DNN models and a configurable mixed OU-based single bit ReRAM DNN engine to enable the algorithm with better performance and energy efficiency. Experimental results show that PQ-PIM achieves up to 1.74x performance improvement, 62% energy saving, and 5.84x compression ratio of occupied crossbars, compared to the state-of-the-art ReRAM-based DNN accelerators.
Keywords:
Resistive random access memory (ReRAM)
Deep Neural Networks (DNNs)
Pruning
Quantization
Accelerator

Journal

Journal of Systems Architecture cover
Journal of Systems Architecture
IF:
4.1
Papers:
2.9K
Citations:
4.2K

Organization

S
shandong university
Scholars:
9.1W
Papers: 6.3W
Citations: 94
Cited Papers

Cited Papers

Citing Papers

Citing Papers