arrow
Return

High Performance Singular Value Decomposition on GPU Architectures

delete2026-03-01
delete0
PRE
AI
H
Hansheng Wang
S
Shaoshuai Zhang *
R
Ruiyi Zhan
W
Wenjing Huang
R
Runzhi Hu
J
Jun Chen
L
Li, Qiao
H
Hancong Duan
T
Tan, Guangming
T
Tao, Dingwen
DOI:10.1145/3787861delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
With the advancement of GPU architecture, matrix computation engines such as NVIDIA Tensor Cores now support double-precision (FP64) General matrix multiplications (GEMMs) with the same efficiency as single-precision (FP32) GEMMs. However, the adoption of this enhanced FP64 capability remains limited, primarily restricted to applications that involve multiple FP64 BLAS3 operations. Singular Value Decomposition (SVD), a fundamental decomposition in numerical linear algebra with numerous applications, can greatly benefit from exploiting this hardware feature. In this article, for FP32 SVD, we propose a novel algorithm, FP64 precision eigenvalue decomposition (EVD) based SVD, specifically designed to leverage the latest GPU architectural features. We provide a theoretical analysis demonstrating the feasibility of our approach on emerging GPU architectures and evaluate it from both accuracy and performance perspectives. Moreover, for FP64 SVD, we introduce a double-blocking band reduction technique combined with a GPU-based bulge chasing algorithm to further accelerate the overall SVD process. Experimental results show that, for FP32 SVD, our EVD-based SVD implementation achieves higher numerical accuracy and delivers speedups of up to 6.1 & times; on H100 and 5.0 & times; on A100 over the state-of-the-art cuSOLVER SVD solver. In the case of FP64 SVD, our method also achieves 4.9 & times; and 4.8 & times; speedups on H100 and A100, respectively. These results highlight the potential of our approach as a highly efficient and accurate solution for SVD on modern GPU platforms.
Keywords:
Singular value decomposition
SVD
eigenvalue decomposition
EVD
bidiagonalization
two-stage
double-blocking
band reduction
bulge chasing
numerical linear algebra

Journal

A
ACM Transactions on Architecture and Code Optimization
IF:
1.8
Papers:
96
Citations:
1.1K

Organization

U
university of electronic science & technology of china
Scholars:
3.2K
Papers: 970
Citations: 0
I
institute of computing technology, cas
Scholars:
1.0K
Papers: 877
Citations: 1
S
sichuan university
Scholars:
11.8W
Papers: 7.7W
Citations: 100
C
chinese academy of sciences
Scholars:
55.9W
Papers: 44.7W
Citations: 704
researcher View more organizations