arrow
Return

Compiler-based loop unrolling optimization for dual-SIMD extensions

delete2026-03-07
delete0
PRE
AI
J
Jinyang Yao
L
Lili Liu
X
Xuanyu Fu
W
Wenbo Liu
C
Chaowei Zhao
W
Wei Wu *
Z
Zheng Shan *
DOI:10.1007/s11227-026-08404-wdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Single Instruction Multiple Data (SIMD) extensions play a pivotal role in accelerating computations for both high-performance computing (HPC) and artificial intelligence (AI) workloads. To further exploit these architectures, this study introduces a compiler-based approach that generates parallel code for processors equipped with dual-SIMD extensions via vectorizable loop unrolling. The proposed optimization is implemented in both the GNU Compiler Collection (GCC) and the SWGCC710 compilers as an optional compilation pass that can be activated through a single flag. Performance evaluations conducted on Shenwei and Intel processors–using the SPEC CPU 2006, SPEC CPU 2017, and NAS Parallel Benchmarks (NPB)–demonstrate the effectiveness of the proposed method. Compared with the conventional-O3 optimization level, the technique achieves a geometric mean speedup of 1.031 and up to 1.165 for full applications, whereas kernel loops attain 1.168 on the Shenwei platform. Experiments conducted on Intel systems further confirm the optimization’s portability and its consistent performance improvements across diverse architectures.
Keywords:
Compiler optimization
Loop unrolling
Dual-SIMD extensions
Pipelining
Autovectorization

Journal

T
The Journal of Supercomputing
IF:
0
Papers:
647
Citations:
0

Organization

Z
zhengzhou university
Scholars:
1.3W
Papers: 3.5K
Citations: 2
I
Information Engineering University
Scholars:
484
Papers: 161
Citations: 0
researcher View more organizations
Cited Papers

Cited Papers

errShare
errSave
ALBUS: A method for efficiently processing SpMV using SIMD and Load balancing
err2021-03-01
err20
PREAI
errBian, Haodong; Huang, Jianqiang; Liu, Lingbin; Huang, Dongqiang; Wang, Xiaoying
errShare
errSave
Partial control-flow linearization
err2018-06-11
err0
PREAI
errSimon Moll; Sebastian Hack
errShare
errSave
The ARM Scalable Vector Extension
err2017-03-01
err0
errOAAI
errNigel Stephens; Stuart Biles; Matthias Boettcher; Jacob Eapen; Mbou Eyole; Giacomo Gabrielli; Matt Horsnell; Grigorios Magklis; Alejandro Martinez; Nathanael Premillieu; Alastair Reid; Alejandro Rico; Paul Walker
errShare
errSave
The Nas Parallel Benchmarks
err1991-09-01
err0
PREAI
errD.H. Bailey; E. Barszcz; J.T. Barton; D.S. Browning; R.L. Carter; L. Dagum; R.A. Fatoohi; P.O. Frederickson; T.A. Lasinski; R.S. Schreiber; H.D. Simon; V. Venkatakrishnan; S.K. Weeratunga
errShare
errSave
no more