arrow
返回

SPCIM: Sparsity-Balanced Practical CIM Accelerator With Optimized Spatial-Temporal Multi-Macro Utilization

delete2023-01-01
delete5
PRE
AI
Y
Yiqi Wang
F
Fengbin Tu
L
Leibo Liu
S
Shaojun Wei
Y
Yuan Xie
S
Shouyi Yin *
DOI:10.1109/TCSI.2022.3216735delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Compute-in-memory (CIM) is a promising technique that reduces data movement in neural network (NN) acceleration. To achieve higher efficiency, some recent CIM accelerators exploit NN sparsity based on CIM's small-grained operation unit (OU) feature. However, new problems arise in a practical multi-macro accelerator: The mismatch between workload parallelism and CIM macro organization causes spatial under-utilization; The multiple macros' different computation time leads to temporal under-utilization. To solve the under-utilization problems, we propose a Sparsity-balanced Practical CIM accelerator (SPCIM), including optimized dataflow and hardware architecture design. For the CIM dataflow design, we first propose a reconfigurable cluster topology for CIM macro organization. Then we regularize weight sparsity in the OU-height pattern and reorder the weight matrix based on the sparsity ratio. The cluster topology can be reshaped to match workload parallelism for higher spatial utilization. Each CIM cluster's workload is dynamically rebalanced for higher temporal utilization. Our hardware architecture supports the proposed dataflow with a spatial input dispatcher and a temporal workload allocator. Experimental results show that, compared with the baseline sparse CIM accelerator that suffers from spatial and temporal under-utilization, SPCIM achieves 2.94 x speedup and 2.86 x energy saving. The proposed sparsity-balanced dataflow and architecture are generic and scalable, which can be applied to other CIM accelerators. We strengthen two state-of-the-art CIM accelerators with the SPCIM techniques, improving their energy efficiency by 1.92 x and 5.59 x , respectively.
Keyword:
Computer architecture
Common Information Model (computing)
Topology
Parallel processing
Organizations
Artificial neural networks
Spatial databases
Compute-in-memory (CIM)
neural network
sparsity
CIM dataflow
CIM accelerator

期刊

IEEE Transactions on Circuits and Systems I-Regular Papers 封面图
IEEE Transactions on Circuits and Systems I-Regular Papers
IF:
5.2
论文数:
9.7K
被引数:
2.2W

机构

T
tsinghua university
学者数:
11.9W
论文数: 10.0W
被引数: 137
University of California System 封面图
University of California System
学者数:
37.5W
论文数: 33.7W
被引数: 6.6K
引用论文

引用论文

Polaron in a Quasi 0D Nanocrystal
err2006-01-13
err0
PREAI
errL C Fai; A J Fotue; V B Mborong; S Domngang; N Issofa; D Tchassem
err分享
err收藏
Retreatment with anti‐PD‐1 antibody in non‐small cell lung cancer patients previously treated with anti‐PD‐L1 antibody
err2019-11-07
err0
errOAAI
errKohei Fujita; Yuki Yamamoto; Osamu Kanai; Misato Okamura; Masayuki Hashimoto; Koichi Nakatani; Satoru Sawai; Tadashi Mio
err分享
err收藏
Persuasion in Parallel
err
IF0
err2022-01-01
err0
PREAI
errAlexander Coppock
err分享
err收藏
Conductivity of zirconium oxide alloyed with lithium oxide
err2008-10-10
err0
PREAI
errN. Yu. Nagaeva; A. A. Surin; L. A. Blaginina; V. P. Obrosov
err分享
err收藏
err分享
err收藏
More is Less: Domain-Specific Speech Recognition Microprocessor Using One-Dimensional Convolutional Recurrent Neural Network
err2022-04-01
err14
PREAI
errLiu, Bo; Cai, Hao; Zhang, Zilong; Ding, Xiaoling; Wang, Ziyu; Gong, Yu; Liu, Weiqiang; Yang, Jinjiang; Wang, Zhen; Yang, Jun
err分享
err收藏
Preoperative expectations for health‐related quality of life after lung transplant
err2018-09-25
err0
PREAI
errMeghan Aversa; Noori A. Chowdhury; George Tomlinson; Lianne G. Singer
err分享
err收藏
err分享
err收藏
学者 查看更多内容