arrow
返回

Optimizing Memory Bandwidth Use and Performance for Matrix-Vector Multiplication in Iterative Methods

delete2011-08-22
delete10
PRE
AI
D
David Boland *
G
George A. Constantinides
DOI:10.1145/2000832.2000834delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Computing the solution to a system of linear equations is a fundamental problem in scientific computing, and its acceleration has drawn wide interest in the FPGA community [Morris et al. 2006; Zhang et al. 2008; Zhuo and Prasanna 2006]. One class of algorithms to solve these systems, iterative methods, has drawn particular interest, with recent literature showing large performance improvements over General-Purpose Processors (GPPs) [Lopes and Constantinides 2008]. In several iterative methods, this performance gain is largely a result of parallelization of the matrix-vector multiplication, an operation that occurs in many applications and hence has also been widely studied on FPGAs [Zhuo and Prasanna 2005; El-Kurdi et al. 2006]. However, whilst the performance of matrix-vector multiplication on FPGAs is generally I/O bound [Zhuo and Prasanna 2005], the nature of iterative methods allows the use of on-chip memory buffers to increase the bandwidth, providing the potential for significantly more parallelism [deLorimier and DeHon 2005]. Unfortunately, existing approaches have generally only either been capable of solving large matrices with limited improvement over GPPs [Zhuo and Prasanna 2005; El-Kurdi et al. 2006; deLorimier and DeHon 2005], or achieve high performance for relatively smallmatrices [Lopes and Constantinides 2008; Boland and Constantinides 2008]. This article proposes hardware designs to take advantage of symmetrical and banded matrix structure, as well as methods to optimize the RAM use, in order to both increase the performance and retain this performance for larger-order matrices.
Keyword:
Algorithms
Design
Performance
Iterative methods
integer linear programming
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

ACM Transactions on Reconfigurable Technology and Systems 封面图
ACM Transactions on Reconfigurable Technology and Systems
IF:
2.8
论文数:
597
被引数:
810

机构

I
Imperial College London
学者数:
8.3W
论文数: 7.3W
被引数: 11.1W
引用论文

引用论文

Novel High-Temperature, High-Power Handling All-Cu Interconnections through Low-Temperature Sintering of Nanocopper Foams
err2016-05-01
err0
PREAI
errNinad Shahane; Kashyap Mohan; Rakesh Behera; Antonia Antoniou; Pulugurtha Raj Markondeya; Vanessa Smet; Rao Tummala
err分享
err收藏
没有更多内容