arrow
Return

Toggle Rate Aware Quantization Model Based on Digital Floating-Point Computing-In-Memory Architecture

delete2024-06-01
delete1
PRE
AI
X
Xi Chen
Y
Yitong Zhao
A
An Guo
J
Jinwu Chen
F
Fangyuan Dong
张朝阳 (Zhaoyang Zhang)
T
Tianzhu Xiong
B
Bo Wang
Y
Yuyao Kong
X
Xin Si *
DOI:10.1109/TCSII.2024.3354313delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Computing-in-memory (CIM) has been proven to achieve high energy efficiency and significant acceleration effects on neural networks with high computational parallelism. Based on typical integer CIMs, some floating-point CIMs (FP-CIM) are proposed recently to execute more accuracy-demanding tasks such as training and high-precision inference. However, prior research has not adequately explored the relationship between circuit design within the FP-CIM architecture and hardware/software metrics. Furthermore, in digital circuits, the data toggle rate significantly affect hardware performance. In this brief, a toggle rate-aware quantization model is proposed to define and explore the design space of FP-CIM. Based on the experimental results, some key considerations on FP-CIM design are derived. With the toggle rate reduction scheme, toggle rate can be reduced by 28%, resulting in a remarkable 1.18x improvement in energy efficiency with only a 0.35% accuracy loss. To validate our model, a 28nm digital FP-CIM test chip is fabricated which achieves energy efficiency of 32.28 TFLOPS/W and inference accuracy of 76.14% on DenseNet161 and ImageNet dataset.
Keywords:
Quantization (signal)
Measurement
Computer architecture
Integrated circuit modeling
Energy efficiency
Computational modeling
Space exploration
Quantization
toggle rate
floating-point
computing-in-memory

Journal

I
IEEE Transactions on Circuits and Systems and Express Briefs
IF:
4.9
Papers:
8.8K
Citations:
2.5W

Organization

S
southeast university - china
Scholars:
5.3W
Papers: 4.9W
Citations: 57