返回
High Performance CNN Accelerators Based on Hardware and Algorithm Co-Optimization
DOI:10.1109/TCSI.2020.3030663.png)
摘要
En 中文
Convolutional neural networks (CNNs) have been widely used in image classification and recognition due to their effectiveness; however, CNNs use a large volume of weight data that is difficult to store in on-chip memory of embedded designs. Pruning can compress the CNN model at a small accuracy loss; however, a pruned CNN model operates slower when implemented on a parallel architecture. In this paper, a hardware-oriented CNN compression strategy is proposed; a deep neural network (DNN) model is divided into no-pruning layers (N P-layers) and pruning layers (P-layers). A NP-layer has a regular weights distribution for parallel computing and high performance. A P-layer is irregular due to pruning, but it generates a high compression ratio. Uniform and incremental quantization schemes are used to achieve a tradeoff between compression ratio and processing efficiency at a small loss in accuracy. A distributed convolutional architecture with several parallel finite impulse response (FIR) filters is further proposed for the regular model in the NP-layers. A shift-accumulator based processing element with an activation-driven data flow (ADF) is proposed for the irregular sparse model in the P-layers. Based on the proposed compression strategy and hardware architecture, a hardware/algorithm co-optimization (HACO) approach is proposed for implementing a NP - P hybrid compressed CNN model on FPGAs. For a hardware accelerator on a single FPGA chip without the use of off-chip memory, a 27.5x compression ratio is achieved with 0.44% top-5 accuracy loss for VGG-16. The implementation of the compressed VGG-16 model on a Xilinx VCU118 evaluation board processes 83.0 frames per second (FPS) for image applications, this is 1.8x superior than the state-of-the-art design found in the technical literature.
Keyword:
Convolutional neural network (CNN)
field programmable gate array (FPGA)
network compression
hardware acceleration
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
5.2
论文数:
9.7K
被引数:
2.2W
机构
引用论文
A Retrospective and Prospective View of Approximate Computing [Point of View}
PROCEEDINGS OF THE IEEE
IF25.9
Do Perceptions of Competence Mediate The Relationship Between Fundamental Motor Skill Proficiency and Physical Activity Levels of Children in Kindergarten?能力的感知是否可以介导幼儿园儿童的基本运动技能熟练程度与身体活动水平之间的关系?
Improved outcome in HLA-identical sibling hematopoietic stem-cell transplantation for acute myelogenous leukemia predicted by KIR and HLA genotypes
Blood
IF0

