arrow
Return

GRASP: Accelerating Hash-Based PQC Performance on GPU Parallel Architecture

delete2026-01-16
delete0
PRE
AI
Y
Yijing Ning
J
Jiankuo Dong
林家元 (Jingqiang Lin)
F
Fangyu Zheng
Y
Y. W. Fu
F
Fu Xiao
DOI:10.1109/TC.2026.3654650delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
SPHINCS<inline-formula><tex-math notation="LaTeX">${}^{+}$</tex-math></inline-formula>, one of the Post-Quantum Cryptography Digital Signature Algorithms (PQC-DSA) selected by NIST in the third round, features very short public and private key lengths but faces significant performance challenges compared to other post-quantum cryptographic schemes, limiting its suitability for real-world applications. In scenarios involving a large number of concurrent signing or verification tasks, these performance bottlenecks become particularly critical. To address these challenges, we propose the <i>G</i>PU-based pa<i>R</i>allel <i>A</i>ccelerated <i>SP</i>HINCS<inline-formula><tex-math notation="LaTeX">${}^{+}$</tex-math></inline-formula> (<i>GRASP</i>), which leverages GPU technology to enhance the efficiency of SPHINCS<inline-formula><tex-math notation="LaTeX">${}^{+}$</tex-math></inline-formula> signing and verification processes. We propose an adaptable parallelization strategy for SPHINCS<inline-formula><tex-math notation="LaTeX">${}^{+}$</tex-math></inline-formula>, analyzing its signing and verification processes to identify critical sections for efficient parallel execution. Utilizing CUDA, we perform bottom-up optimizations, focusing on memory access patterns and hypertree computation, to enhance GPU resource utilization. These efforts, combined with kernel fusion technology, result in significant improvements in throughput and overall performance. Compared to previous works, our approach achieves the highest occupancy. Extensive experimentation demonstrates that our optimized CUDA implementation of SPHINCS<inline-formula><tex-math notation="LaTeX">${}^{+}$</tex-math></inline-formula> achieves superior performance. Specifically, our GRASP scheme delivers throughput improvements ranging from 1.09× to 3.45× compared to state-of-the-art GPU-based solutions and surpasses the NIST reference implementation by over three orders of magnitude, highlighting a significant performance advantage.
Keywords:
PQC
hash-based digital signature
SPHINCS ${}^{+}$ +
GPU
CUDA

Journal

IEEE Transactions on Computers cover
IEEE Transactions on Computers
IF:
3.8
Papers:
5.3K
Citations:
9.8K

Organization

U
university of chinese academy of sciences
Scholars:
976
Papers: 395
Citations: 0
N
Nanjing University of Posts and Telecommunications
Scholars:
2.4K
Papers: 969
Citations: 1.2W
U
University of Science and Technology of China
Scholars:
1.6W
Papers: 5.5K
Citations: 11.3W
researcher View more organizations