arrow
返回

STuning-DL: Model-Driven Autotuning of Sparse GPU Kernels for Deep Learning

delete2024-01-01
delete0
delete
OA
AI
R
Roberto L. Castro *
D
Diego Andrade
B
Basilio B. Fraguela
DOI:10.1109/ACCESS.2024.3402326delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The relentless growth of modern Machine Learning models has spurred the adoption of sparsification techniques to simplify their architectures and reduce the computational demands. Network pruning has demonstrated success in maintaining original network accuracy while shedding significant portions of the original weights. However, leveraging this sparsity efficiently remains challenging due to computational irregularities, particularly in GPU kernels. A new trend of template-based GPU kernels for semi-structured sparsity shows promise in efficiency but lacks autotuning capabilities to adapt to input dynamics, often underperforming in scenarios where they have not been meticulously hand-tuned. We present STuning-DL, the first pruning-aware autotuner for third-party template-based implementations enabling efficient optimization of sparse kernels for Deep Learning, spanning from high-level aspects (CUDA C++ level) down to GPU-native instructions specifics (assembly-level). STuning-DL tunes and optimizes at run-time sparse kernels' performance for each input problem, yielding speedups of up to 5.42 x on NVIDIA T4-16GB and up to 3.6 x on NVIDIA A100-40GB GPU in sparse matrices from real world models compared to existing heuristics from sparse libraries like cuSparse and cuSparseLt.
Keyword:
CUDA
GPU
learning-based predictive model
network pruning
sparse computation
SpMM
Tensor Core

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

U
Universidade da Coruna
学者数:
6.6K
论文数: 5.7K
被引数: 11
引用论文

引用论文

NAP: Neural architecture search with pruningNAP: 神经架构搜索与修剪
err2022-03-01
err22
PREAI
errDing, Yadong; Wu, Yu; Huang, Chengyue; Tang, Siliang; Wu, Fei; Yang, Yi; Zhu, Wenwu; Zhuang, Yueting
err分享
err收藏
A Survey on Compiler Autotuning using Machine Learning基于机器学习的编译器自动调优研究综述
err2018-09-18
err123
errOAAI
errAshouri, Amir H.; Killian, William; Cavazos, John; Palermo, Gianluca; Silvano, Cristina
err分享
err收藏
Prognostic Significance of Perineural Invasion in Cervical Cancer
err2013-03-01
err0
PREAI
errHyun Chul Cho; Haeryoung Kim; Hye-yon Cho; Kidong Kim; Jae Hong No; Yong-Beom Kim
err分享
err收藏
Transcription Impacts the Efficiency of mRNA Translation via Co-transcriptional N6-adenosine Methylation
errCell
IF0
err2017-04-01
err0
errOAAI
errBoris Slobodin; Ruiqi Han; Vittorio Calderone; Joachim A.F. Oude Vrielink; Fabricio Loayza-Puch; Ran Elkon; Reuven Agami
err分享
err收藏
Genetic structure of bisexual and parthenogenetic populations of Artemia from Italian brackish–hypersaline waters
err2003-03-01
err0
errOAAI
errGiuseppe Nascetti; Paola Bondanelli; Antonella Aldinucci; Roberta Cimmaruta
err分享
err收藏
err分享
err收藏
Therapeutic potential of cinnamon for neurological disorders: A mini-review
err2022-03-01
err0
errOAAI
errAli Ahmadi; Mahdyieh Naziri; Fatemeh Fallahpour; Kosar Gholami; Javad Arabpour; Fateme Pazeshgare; Diba Akbarzadeh; Arina Ansari; Hamoun Sabri; Niloofar Deravi
err分享
err收藏
err分享
err收藏
学者 查看更多内容