arrow
Return

SqueezeFlow: A Sparse CNN Accelerator Exploiting Concise Convolution Rules

delete2019-11-01
delete34
PRE
AI
J
Jiajun Li
S
Shuhao Jiang
S
Shijun Gong
J
Jingya Wu
J
Junchao Yan
G
Guihai Yan
X
Xiaowei Li *
DOI:10.1109/TC.2019.2924215delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Convolutional Neural Networks (CNNs) have been widely used in machine learning tasks. While delivering state-of-the-art accuracy, CNNs are known as both compute- and memory-intensive. This paper presents the SqueezeFlow accelerator architecture that exploits sparsity of CNN models for increased efficiency. Unlike prior accelerators that trade complexity for flexibility, SqueezeFlow exploits concise convolution rules to benefit from the reduction of computation and memory accesses as well as the acceleration of existing dense architectures without intrusive PE modifications. Specifically, SqueezeFlow employs a PT-OS-sparse dataflow that removes the ineffective computations while maintaining the regularity of CNN computations. We present a full design down to the layout at 65 nm, with an area of 4.80mm2 and power of 536.09mW. The experiments show that SqueezeFlow achieves a speedup of 2:9 on VGG16 compared to the dense architectures, with an area and power overhead of only 8.8 and 15.3 percent, respectively. On three representative sparse CNNs, SqueezeFlow improves the performance and energy efficiency by 1:8 and 1:5 over the state-of-the-art sparse accelerators.
Keywords:
Convolutional neural networks
accelerator architecture
hardware acceleration
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Computers cover
IEEE Transactions on Computers
IF:
3.8
Papers:
5.3K
Citations:
9.8K

Organization

C
chinese academy of sciences
Scholars:
56.0W
Papers: 44.8W
Citations: 704