Return
Toggle Rate Aware Quantization Model Based on Digital Floating-Point Computing-In-Memory Architecture
DOI:10.1109/TCSII.2024.3354313.png)
Abstract
En 中文
Computing-in-memory (CIM) has been proven to achieve high energy efficiency and significant acceleration effects on neural networks with high computational parallelism. Based on typical integer CIMs, some floating-point CIMs (FP-CIM) are proposed recently to execute more accuracy-demanding tasks such as training and high-precision inference. However, prior research has not adequately explored the relationship between circuit design within the FP-CIM architecture and hardware/software metrics. Furthermore, in digital circuits, the data toggle rate significantly affect hardware performance. In this brief, a toggle rate-aware quantization model is proposed to define and explore the design space of FP-CIM. Based on the experimental results, some key considerations on FP-CIM design are derived. With the toggle rate reduction scheme, toggle rate can be reduced by 28%, resulting in a remarkable 1.18x improvement in energy efficiency with only a 0.35% accuracy loss. To validate our model, a 28nm digital FP-CIM test chip is fabricated which achieves energy efficiency of 32.28 TFLOPS/W and inference accuracy of 76.14% on DenseNet161 and ImageNet dataset.
Keywords:
Quantization (signal)
Measurement
Computer architecture
Integrated circuit modeling
Energy efficiency
Computational modeling
Space exploration
Quantization
toggle rate
floating-point
computing-in-memory
Journal
I
IF:
4.9
Papers:
8.8K
Citations:
2.5W

