arrow
Return

Performance analysis and optimization for SpMV based on aligned storage formats on an ARM processor

delete2021-12-01
delete12
PRE
AI
Y
Yufeng Zhang
W
Wangdong Yang *
李肯立 cover
李肯立 (Kenli Li)
D
Dahai Tang
李克勤 cover
李克勤 (Keqin Li)
DOI:10.1016/j.jpdc.2021.08.002delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Sparse matrix-vector multiplication (SpMV) has always been a hot topic of research for scientific computing and big data processing, but the sparsity and discontinuity of the nonzero elements in a sparse matrix lead to the memory bottleneck of SpMV. In this paper, we propose aligned CSR (ACSR) and aligned ELL (AELL) formats and a parallel SpMV algorithm to utilize NEON SIMD registers on ARM processors. We analyze the impact of SIMD instruction latency, cache access, and cache misses on SpMV with different formats. In the experiments, our SpMV algorithm based on ACSR achieves 1.18x and 1.56x speedup over SpMV based on CSR and SpMV in PETSc, respectively, and AELL achieves 1.21x speedup over ELL. The deviations between the theoretical results and experimental results in the instruction latency and cache access are 10.26% and 10.51% in ACSR and 5.68% and 2.91% in AELL, respectively. (C) 2021 Elsevier Inc. All rights reserved.
Keywords:
ARM
NEON
SIMD
SpMV
Storage formats

Journal

Journal of Parallel and Distributed Computing cover
Journal of Parallel and Distributed Computing
IF:
4
Papers:
3.8K
Citations:
4.8K

Organization

H
hunan university
Scholars:
4.5W
Papers: 3.3W
Citations: 70