Return
PQ-PIM: A pruning-quantization joint optimization framework for ReRAM-based-in DNN accelerator
Y
X
X
Y
Z
Z
DOI:10.1016/j.sysarc.2022.102531.png)
Abstract
En 中文
Pruning and quantization are two efficient techniques to achieve performance improvement and energy saving for ReRAM-based DNN accelerators. However, most existing ReRAM-based DNN accelerators using pruning and quantization are based on an overidealized multi-bit ReRAM crossbar while neglecting the practical structure constraints. Due to the restriction of immature process technology, the actual matrix-vector multiplication must be conducted in a smaller operation unit (OU) granularity with single bit ReRAM cells. In this paper, we propose an efficient pruning-quantization joint exploration framework for practical ReRAM-based DNN accelerator, termed as PQ-PIM, which consists of a patch-wise pruning-quantization algorithm based on patch importance analysis to compress DNN models and a configurable mixed OU-based single bit ReRAM DNN engine to enable the algorithm with better performance and energy efficiency. Experimental results show that PQ-PIM achieves up to 1.74x performance improvement, 62% energy saving, and 5.84x compression ratio of occupied crossbars, compared to the state-of-the-art ReRAM-based DNN accelerators.
Keywords:
Resistive random access memory (ReRAM)
Deep Neural Networks (DNNs)
Pruning
Quantization
Accelerator
Journal
IF:
4.1
Papers:
2.9K
Citations:
4.2K
