arrow
返回

Runtime Programmable and Memory Bandwidth Optimized FPGA-Based Coprocessor for Deep Convolutional Neural Network

delete2018-12-01
delete35
PRE
AI
N
Nimish Shah
P
Paragkumar Chaudhari
K
Kuruvilla Varghese *
DOI:10.1109/TNNLS.2018.2815085delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The deep convolutional neural network (DCNN) is a class of machine learning algorithms based on feed-forward artificial neural network and is widely used for image processing applications. Implementation of DCNN in real-world problems needs high computational power and high memory bandwidth, in a power-constrained environment. A general purpose CPU cannot exploit different parallelisms offered by these algorithms and hence is slow and energy inefficient for practical use. We propose a field-programmable gate array (FPGA)-based runtime programmable coprocessor to accelerate feed-forward computation of DCNNs. The coprocessor can be programmed for a new network architecture at runtime without resynthesizing the FPGA hardware. Hence, it acts as a plug-and-use peripheral for the host computer. Caching is implemented for input features and filter weights using on-chip memory to reduce the external memory bandwidth requirement. Data are prefetched at several stages to avoid stalling of computational units and different optimization techniques are used to efficiently reuse the fetched data. Dataflow is dynamically adjusted in runtime for each DCNN layer to achieve consistent computational throughput across a wide range of input feature sizes and filter sizes. The coprocessor is prototyped using Xilinx Virtex-7 XC7VX485T FPGA-based VC707 board and operates at 150 MHz. Experimental results show that our implementation is 15x energy efficient than highly optimized CPU implementation and achieves consistent computational throughput of more than 140 G operations/s for a wide range of input feature sizes and filter sizes. Off-chip memory transactions decrease by 111x due to the use of the on-chip cache.
Keyword:
Accelerator
coprocessor
deep convolutional neural network (DCNN)
deep learning
field-programmable gate array (FPGA)
runtime programmable
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Neural Networks and Learning Systems 封面图
IEEE Transactions on Neural Networks and Learning Systems
IF:
8.9
论文数:
7.6K
被引数:
7.2W

机构

I
indian institute of science (iisc) - bangalore
学者数:
1.4W
论文数: 1.4W
被引数: 11
引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
Membrane potential correlates of sensory perception in mouse barrel cortex
err2013-10-06
err0
PREAI
errShankar Sachidhanandam; Varun Sreenivasan; Alexandros Kyriakatos; Yves Kremer; Carl C H Petersen
err分享
err收藏
Field evaluation and risk assessment of transgenic indica basmati rice
err2004-05-01
err0
PREAI
errKhurram Bashir; Tayyab Husnain; Tahira Fatima; Zakia Latif; Syed Aks Mehdi; Sheikh Riazuddin
err分享
err收藏
Perceptually‐motivated Real‐time Temporal Upsampling of 3D Content for High‐refresh‐rate Displays
err2010-06-07
err0
errOAAI
errPiotr Didyk; Elmar Eisemann; Tobias Ritschel; Karol Myszkowski; Hans‐Peter Seidel
err分享
err收藏
Generation of Allospecific Natural Killer Cells by Stimulation Across a Polymorphism of HLA-C
err1993-05-21
err0
PREAI
errMarco Colonna; Edward G. Brooks; Michela Falco; Giovan Battista Ferrara; Jack L. Strominger
err分享
err收藏
DANoC: An Efficient Algorithm and Hardware Codesign of Deep Neural Networks on Chip
err2017-01-01
err10
PREAI
errZhou, Xichuan; Li, Shengli; Tang, Fang; Hu, Shengdong; Lin, Zhi; Zhang, Lei
err分享
err收藏
Hierarchical Address Event Routing for Reconfigurable Large-Scale Neuromorphic Systems
err2017-10-01
err84
errOAAI
errPark, Jongkil; Yu, Theodore; Joshi, Siddharth; Maier, Christoph; Cauwenberghs, Gert
err分享
err收藏
学者 查看更多内容