arrow
Return

TIMAQ: A Time-Domain Computing-in-Memory-Based Processor Using Predictable Decomposed Convolution for Arbitrary Quantized DNNs

delete2021-10-01
delete16
PRE
AI
J
Jianxun Yang
Y
Yuyao Kong
Z
Zhao Zhang
Z
Zhuangzhi Liu
J
Jing Zhou
Y
Yiqi Wang
刘永刚 (Yonggang Liu)
C
Chenfu Guo
T
Te Hu
C
Congcong Li
L
Leibo Liu
J
Jin Zhang
S
Shaojun Wei
J
Jun Yang
S
Shouyi Yin *
DOI:10.1109/JSSC.2021.3095232delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Energy-efficient processors are crucial for accelerating deep neural networks (DNNs) on edge devices with limited battery capacity. To reduce energy consumption, time-domain computing-in-memory (TD-CIM) is a splendid architecture, which consumes low computation and memory access energy due to low toggle rate of time-based signals and less data movements, respectively. When deploying DNNs in TD-CIMs, quantization is required, which has two types: uniform quantization (UQ) and nonuniform quantization (NUQ). To reach the same accuracy for one DNN, NUQ achieves smaller model size than UQ. Due to varying weight distributions across layers, mixed-precision quantization can further reduce model size, without degrading accuracy. However, previous TD-CIMs are inefficient for mixed-precision NUQ-DNNs due to their adopted bit-serial convolution increasing computation amount significantly. To address that, we propose a unique-weight convolution to accelerate mixed-precision NUQ-DNNs by a special kernel decomposition, reducing computation count remarkably. Based on that, we design a TD-CIM-based processor, TIMAQ, with three architectural techniques: 1) bit-cross-flipping-based kernel decomposer to reduce memory accesses and operations of decomposing kernels; 2) dual-mode-complementary predictor to remove redundant computations; and 3) activation-weight-adaptive pulse quantizer to decrease pulse quantization energy and error. Fabricated in 28-nm CMOS technology and tested on 1-8-b NUQ-DNNs, TIMAQ achieves 2.4-152.7-TOPS/W peak energy efficiency.
Keywords:
Indexes
Kernel
Convolution
Quantization (signal)
Time-domain analysis
Memory management
Delays
Computing-in-memory (CIM)
deep neural network (DNN) processor
nonuniform quantization (NUQ)
redundant computation
time-domain (TD) computing

Journal

I
IEEE Journal of Solid-State Circuits
IF:
5.6
Papers:
888
Citations:
2.7W

Organization

T
tsinghua university
Scholars:
11.8W
Papers: 10.0W
Citations: 137
S
southeast university - china
Scholars:
5.3W
Papers: 4.9W
Citations: 57