Return
Efficient sparse matrix-vector multiplication using cache oblivious extension quadtree storage format
DOI:10.1016/j.future.2015.03.005.png)
Abstract
En 中文
In this paper, we elaborate on improving the sparse matrix storage format to optimize the data locality of sparse matrix-vector multiplication (SpMVM) algorithm, and its parallel performance. First of all, we propose a cache oblivious extension quadtree storage structure (COEQT), in which the sparse matrix is recursively divided into sub-regions that can well fit into cache to improve the data locality. Later on, we present a COEQT based SpMVM algorithm and optimize its performance through manual vectorization. With this storage format, the original SpMVM is divided into computations of relatively independent small matrices. In addition, this region-based computation framework is also suitable for high performance computing in distributed computing environment. So, we finally present a parallel SpMVM algorithm based on the proposed COEQT. Extensive and comprehensive experiments show that the sparse matrix-vector multiplication using the COEQT storage format achieves on average 1.1-1.5x speedup compared with CSR format and further higher performance through instruction level optimization techniques. The experiment in Lenovo Deepcomp 7000 demonstrates that this method achieves on average 1.63x speedup compared with the Intel Cluster Math Kernel Library implementation. (C) 2015 Elsevier B.V. All rights reserved.
Keywords:
Sparse matrix-vector multiplication
Sparse matrix storage
Data locality
Cache oblivious
Extension quadtree
Distributed parallelism
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
F
IF:
6.1
Papers:
6.9K
Citations:
2.3W
Organization
Cited Papers
Permeability of Neotropical agricultural lands to a key native ungulate—Are well‐connected forests important?
Biotropica
IF0

