arrow
返回

Enabling Highly Efficient Capsule Networks Processing Through Software-Hardware Co-Design

delete2021-04-01
delete7
PRE
AI
X
Xingyao Zhang *
X
Xin Fu
D
Donglin Zhuang
C
Chenhao Xie
S
Shuaiwen Leon Song
DOI:10.1109/TC.2021.3056929delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
As the demand for the image processing increases, the image features become increasingly complicated. Although the Convolutional Neural Network (CNN) have been widely adopted for the imaging processing tasks, it has been found easily misled due to the massive usage of pooling operations. A novel neural network structure called Capsule Networks (CapsNet) is proposed to address the CNN challenge and essentially enhance the learning ability for the image segmentation and object detection. Since the CapsNet contains the high volume of the matrix execution, it has been generally accelerated on modern GPU platforms with the highly optimized deep-learning library. However, the routing procedure of CapsNet introduces the special program and execution features,including massive unshareable intermediate variables and intensive synchronizations, causing inefficient CapsNet execution on modern GPU. To address these challenges, we propose the software-hardware co-designed optimizations, SH-CapsNet, which includes the software-level optimizations named S-CapsNet and a hybrid computing architecture design named PIM-CapsNet. In software-level, S-CapsNet reduces the computation and memory accesses by exploiting the computational redundancy and data similarity of the routing procedure. In hardware-level, the PIM-CapsNet leverages the processing-in-memory capability of today's 3D stacked memory to conduct the off-chip in-memory acceleration solution for the routing procedure, while pipelining with the GPU's on-chip computing capability for accelerating CNN types of layers in CapsNet. Evaluation results demonstrate that either our software or hardware optimizations can significantly improve the CapsNet execution efficiency. Together, our co-design can achieve greatly improvement on both performance ($3.41\times$3.41x) and energy savings (68.72 percent) for CapsNet inference, with negligible accuracy loss.
Keyword:
Accelerators
domain-specific architectures
machine learning
emerging technologies
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Computers 封面图
IEEE Transactions on Computers
IF:
3.8
论文数:
5.3K
被引数:
9.8K

机构

U
University of Sydney
学者数:
6.5W
论文数: 6.2W
被引数: 90
U
university of houston system
学者数:
1.4W
论文数: 1.4W
被引数: 16
U
university of houston
学者数:
9.7K
论文数: 7.9K
被引数: 11
学者 查看更多机构
引用论文

引用论文

Genetic testing in Neurology
err2020-08-01
err0
PREAI
errHenrietta Lefroy; Victoria Harrison; Andrea H. Németh
err分享
err收藏
学者 查看更多内容