arrow
返回

A MEMORY EFFICIENT AND FAST SPARSE MATRIX VECTOR PRODUCT ON A GPU

delete2011-01-01
delete64
delete
OA
AI
A
Adam Dziekonski *
A
Adam Lamęcki
M
Michał Mrozowski
DOI:10.2528/PIER11031607delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
This paper proposes a new sparse matrix storage format which allows an efficient implementation of a sparse matrix vector product on a Fermi Graphics Processing Unit (GPU). Unlike previous formats it has both low memory footprint and good throughput. The new format, which we call Sliced ELLR-T has been designed specifically for accelerating the iterative solution of a large sparse and complex-valued system of linear equations arising in computational electromagnetics. Numerical tests have shown that the performance of the new implementation reaches 69 GFLOPS in complex single precision arithmetic. Compared to the optimized six core Central Processing Unit (CPU) (Intel Xeon 5680) this performance implies a speedup by a factor of six. In terms of speed the new format is as fast as the best format published so far and at the same time it does not introduce redundant zero elements which have to be stored to ensure fast memory access. Compared to previously published solutions, significantly larger problems can be handled using low cost commodity GPUs with limited amount of on-board memory.
Keyword:
FINITE-ELEMENT-METHOD
FDTD METHOD
SCATTERING
ALGORITHM
UNITS
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

P
Progress in Electromagnetics Research-PIER
IF:
9.3
论文数:
2.9K
被引数:
2.7K

机构

F
fahrenheit universities
学者数:
1.6W
论文数: 1.3W
被引数: 21
引用论文

引用论文

err分享
err收藏
Collagenases and gelatinases and their inhibitors as anticancer agents
err2020-01-01
err0
PREAI
errNilanjan Adhikari; Sk. Abdul Amin; Tarun Jha
err分享
err收藏
学者 查看更多内容