arrow
返回

Rethinking Embedded Blocks for Machine Learning Applications

delete2021-11-30
delete0
PRE
AI
S
S A Rasoulinezhad *
E
Esther Roorda
W
Wilton, Steve
P
Philip H. W. Leong
D
David Boland
DOI:10.1145/3491234delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The underlying goal of FPGA architecture research is to devise flexible substrates that implement a wide variety of circuits efficiently. Contemporary FPGA architectures have been optimized to support networking, signal processing, and image processing applications through high-precision digital signal processing (DSP) blocks. The recent emergence of machine learning has created a new set of demands characterized by: (1) higher computational density and (2) low precision arithmetic requirements. With the goal of exploring this new design space in a methodical manner, we first propose a problem formulation involving computing nested loops over multiply-accumulate (MAC) operations, which covers many basic linear algebra primitives and standard deep neural network (DNN) kernels. A quantitative methodology for deriving efficient coarse-grained compute block architectures from benchmarks is then proposed together with a family of new embedded blocks, called MLBlocks. An MLBlock instance includes several multiply-accumulate units connected via a flexible routing, where each configuration performs a few parallel dot-products in a systolic array fashion. This architecture is parameterized with support for different data movements, reuse, and precisions, utilizing a columnar arrangement that is compatible with existing FPGA architectures. On synthetic benchmarks, we demonstrate that for 8-bit arithmetic, MLBlocks offer 6x improved performance over the commercial Xilinx DSP48E2 architecture with smaller area and delay; and for timemultiplexed 16-bit arithmetic, achieves 2x higher performance per area with the same area and frequency. All source codes and data, along with documents to reproduce all the results in this article, are available at http://github.com/raminrasoulinezhad/MLBlocks.
Keyword:
FPGA Architectures
coarse-grained compute blocks
reconfigurable architecture
neural networks
digital signal processing

期刊

ACM Transactions on Reconfigurable Technology and Systems 封面图
ACM Transactions on Reconfigurable Technology and Systems
IF:
2.8
论文数:
597
被引数:
810

机构

U
University of Sydney
学者数:
6.5W
论文数: 6.2W
被引数: 90
U
University of British Columbia
学者数:
7.0W
论文数: 6.1W
被引数: 8.6W
引用论文

引用论文

FPGA Logic Block Architectures for Efficient Deep Learning Inference
err2020-06-03
err15
PREAI
errEldafrawy, Mohamed; Boutros, Andrew; Yazdanshenas, Sadegh; Betz, Vaughn
err分享
err收藏
Algorithm 64: Quicksort
err1961-07-01
err0
PREAI
errC. A. R. Hoare
err分享
err收藏
Synthesis and Antidepressant Activity of 8-Amino-Substituted 1-Butyl-3-Methylxanthines Containing a Thietane Ring
err2020-02-22
err0
PREAI
errYu. V. Shabalina; F. A. Khaliullin; I. L. Nikitina; A. F. Miftakhova; R. M. Sharafutdinov
err分享
err收藏
Nouveaux jouets : ce que les enfants identifient comme " jouets de garçons " et " jouets de filles "
err2006-01-01
err0
PREAI
errIsabelle D. Cherney; Hilary J. Harper; Jordan A. Winter
err分享
err收藏
Deep Neural Network Approximation for Custom Hardware: Where We've Been, Where We're Going
err2019-05-30
err130
errOAAI
errWang, Erwei; Davis, James J.; Zhao, Ruizhe; Ng, Ho-Cheung; Niu, Xinyu; Luk, Wayne; Cheung, Peter Y. K.; Constantinides, George A.
err分享
err收藏
Perceived Motivating Factors and Barriers for the Completion of Postgraduate Training Among American Pharmacy Students Prior to Beginning Advanced Pharmacy Practice Experiences美国药学学生在开始高级药学实践经历之前,完成研究生培训的感知激励因素与障碍因素
err2017-06-01
err0
errOAAI
errDrayton A. Hammond; Douglas R. Oyler; John W. Devlin; Jacob T. Painter; Scott Bolesta; Joseph M. Swanson; Brett J. Bailey; Trisha Branan; Jeffrey F. Barletta; Brianne Dunn; Jason S. Haney; Paul Juang; Sandra L. Kane-Gill; Tyree H. Kiser; Hira Shafeeq; Debra Skaar; Pamela Smithburger; Jodi Taylor
err分享
err收藏
Al–Th intermetallic compounds. I
err1955-02-10
err0
errOAAI
errP. B. Braun; J. H. van Vucht
err分享
err收藏
学者 查看更多内容