arrow
Return

Starlight: A kernel optimizer for GPU processing

delete2024-05-01
delete1
delete
OA
AI
A
Alberto Zeni *
E
Emanuele Del Sozzo
E
Eleonora D’Arnese
D
Davide Conficconi
M
Marco D. Santambrogio
DOI:10.1016/j.jpdc.2023.104832delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Over the past few years, GPUs have found widespread adoption in many scientific domains, offering notable performance and energy efficiency advantages compared to CPUs. However, optimizing GPU high-performance kernels poses challenges given the complexities of GPU architectures and programming models. Moreover, current GPU development tools provide few high-level suggestions and overlook the underlying hardware. Here we present Starlight, an open-source, highly flexible tool for enhancing GPU kernel analysis and optimization. Starlight autonomously describes Roofline Models, examines performance metrics, and correlates these insights with GPU architectural bottlenecks. Additionally, Starlight predicts potential performance enhancements before altering the source code. We demonstrate its efficacy by applying it to literature genomics and physics applications, attaining speedups from 1.1x to 2.5x over state-of-the-art baselines. Furthermore, Starlight supports the development of new GPU kernels, which we exemplify through an image processing application, showing speedups of 12.7x and 140x when compared against state-of-the-art FPGA- and GPU-based solutions.
Keywords:
Performance analysis
Performance optimization
High performance computing
GPU
Roofline model
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Journal of Parallel and Distributed Computing cover
Journal of Parallel and Distributed Computing
IF:
4
Papers:
3.8K
Citations:
4.8K

Organization

P
Polytechnic University of Milan
Scholars:
2.0W
Papers: 1.8W
Citations: 24