arrow
返回

Evaluating execution time predictions on GPU kernels using an analytical model and machine learning techniques

delete2023-01-01
delete4
PRE
AI
M
Marcos Amarís *
R
Raphael Y. de Camargo
D
Daniel Cordeiro
A
Alfredo Goldman
D
Denis Trystram
DOI:10.1016/j.jpdc.2022.09.002delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Predicting the performance of applications executed on GPUs is a great challenge and is essential for efficient job schedulers. There are different approaches to do this, namely analytical modeling and machine learning (ML) techniques. Machine learning requires large training sets and reliable features, nevertheless it can capture the interactions between architecture and software without manual intervention. In this paper, we compared a BSP-based analytical model to predict the time of execution of kernels executed over GPUs. The comparison was made using three different ML techniques. The analytical model is based on the number of computations and memory accesses of the GPU, with additional information on cache usage obtained from profiling. The ML techniques Linear Regression, Support Vector Machine, and Random Forest were evaluated over two scenarios: first, data input or features for ML techniques were the same as the analytical model and, second, using a process of feature extraction, which used correlation analysis and hierarchical clustering. Our experiments were conducted with 20 CUDA kernels, 11 of which belonged to 6 real-world applications of the Rodinia benchmark suite, and the other were classical matrix-vector applications commonly used for benchmarking. We collected data over 9 NVIDIA GPUs in different machines. We show that the analytical model performs better at predicting when applications scale regularly. For the analytical model a single parameter lambda is capable of adjusting the predictions, minimizing the complex analysis in the applications. We show also that ML techniques obtained high accuracy when a process of feature extraction is implemented. Sets of 5 and 10 features were tested in two different ways, for unknown GPUs and for unknown Kernels. For ML experiments with a process of feature extractions, we got errors around 1.54% and 2.71%, for unknown GPUs and for unknown Kernels, respectively. (c) 2022 Elsevier Inc. All rights reserved.
Keyword:
Performance prediction
Machine learning
GPU applications
Bulk synchronous parallel model

期刊

Journal of Parallel and Distributed Computing 封面图
Journal of Parallel and Distributed Computing
IF:
4
论文数:
3.8K
被引数:
4.8K

机构

C
communaute universite grenoble alpes
学者数:
3.5W
论文数: 2.7W
被引数: 29
U
universidade federal do abc (ufabc)
学者数:
3.5K
论文数: 3.3K
被引数: 1
I
Inria
学者数:
3.5K
论文数: 2.5K
被引数: 343
U
universidade de sao paulo
学者数:
10.6W
论文数: 6.7W
被引数: 93
学者 查看更多机构
引用论文

引用论文

err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
The Enigmatic HOX Genes: Can We Crack Their Code?
err2019-03-07
err0
errOAAI
errZhifei Luo; Suhn K. Rhie; Peggy J. Farnham
err分享
err收藏
A Performance Model for GPUs with Caches
err2015-07-01
err29
errOAAI
errThanh Tuan Dao; Kim, Jungwon; Seo, Sangmin; Egger, Bernhard; Lee, Jaejin
err分享
err收藏
Archaeological sites as Distributed Long-term Observing Networks of the Past (DONOP)
err2020-05-01
err0
errOAAI
errGeorge Hambrecht; Cecilia Anderung; Seth Brewington; Andrew Dugmore; Ragnar Edvardsson; Francis Feeley; Kevin Gibbons; Ramona Harrison; Megan Hicks; Rowan Jackson; Guðbjörg Ásta Ólafsdóttir; Marcy Rockman; Konrad Smiarowski; Richard Streeter; Vicki Szabo; Thomas McGovern
err分享
err收藏
没有更多内容