arrow
Return

Automatic Tuning based on Hardware Performance Counters and Machine Learning

delete2025-12-30
delete0
delete
OA
AI
S
Suren Harutyunynan Gevorgyan
E
Eduardo César
A
Anna Sikora
J
Jiří Filipovič
J
Jordi Alcaraz
DOI:10.1016/j.future.2025.108358delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
This paper presents a Machine Learning (ML) methodology for automatically tuning parallel applications in heterogeneous High Performance Computing (HPC) environments using Hardware Performance Counters (HwPCs). The methodology addresses three critical challenges: counter quantity versus accessibility tradeoff, data interpretation complexity, and dynamic optimization needs. The introduced ensemble-based methodology automatically identifies minimal yet informative HwPC sets for code region identification and tuning parameter optimization. Experimental validation demonstrates high accuracy in predicting optimal thread allocation ( > 0.90 K-fold accuracy) and thread affinity ( > 0.95 accuracy) while requiring only 4-6 HwPCs. Compared to search-based methods like OpenTuner, the methodology achieves competitive performance with dramatically reduced optimization time. The architecture-agnostic design enables consistent performance across CPU and GPU platforms. These results establish a foundation for efficient, portable, automatic, and scalable tuning of parallel applications.
Keywords:
Hardware Performance Counters
Automatic Dimension Reduction
Machine Learning Ensembles
Tuning Parameter Optimization
Parallel Region Classification
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

F
Future Generation Computer Systems
IF:
0
Papers:
642
Citations:
0

Organization

No organization information available