arrow
Return

Performance-Aligned LLMs for Generating Fast HPC Code

delete2026-03-18
delete0
PRE
AI
D
Daniel Nichols
P
Pranav Polasam
H
Harshitha Menon
A
Aniruddha Marathe
T
Todd Gamblin
A
Abhinav Bhatelé
DOI:10.1109/TPDS.2026.3675550delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Optimizing scientific software is a difficult task because codebases are often large and complex, and performance can depend upon several factors including the algorithm, its implementation, and hardware among others. Causes of poor performance can originate from disparate sources and be difficult to diagnose. Recent years have seen a multitude of work that use large language models (LLMs) to assist in software development tasks. However, these tools are trained to model the distribution of code as text, and are not specifically designed to understand performance aspects of code. In this work, we introduce a reinforcement learning based methodology to align the outputs of code LLMs with performance. This allows us to build upon the current code modeling capabilities of LLMs and extend them to generate better performing code. We demonstrate that our fine-tuned model improves the expected speedup of generated code over base models for a set of benchmark tasks from 0.9 to 1.6 for serial code and 1.9 to 4.5 for OpenMP parallel code.
Keywords:
Large language models (LLMs)
code generation
performance optimization
reinforcement learning

Journal

IEEE Transactions on Parallel and Distributed Systems cover
IEEE Transactions on Parallel and Distributed Systems
IF:
6
Papers:
5.2K
Citations:
1.1W

Organization

L
lawrence livermore national laboratory
Scholars:
653
Papers: 234
Citations: 0
U
university of maryland
Scholars:
4.4K
Papers: 2.1K
Citations: 1
Cited Papers

Cited Papers

errShare
errSave
errShare
errSave
errShare
errSave
Competition-level code generation with AlphaCode
err2022-12-09
err0
errOAAI
errYujia Li; David Choi; Junyoung Chung; Nate Kushman; Julian Schrittwieser; Rémi Leblond; Tom Eccles; James Keeling; Felix Gimeno; Agustin Dal Lago; Thomas Hubert; Peter Choy; Cyprien de Masson d’Autume; Igor Babuschkin; Xinyun Chen; Po-Sen Huang; Johannes Welbl; Sven Gowal; Alexey Cherepanov; James Molloy; Daniel J. Mankowitz; Esme Sutherland Robson; Pushmeet Kohli; Nando de Freitas; Koray Kavukcuoglu; Oriol Vinyals
errShare
errSave
LM4HPC: Towards Effective Language Model Application in High-Performance Computing
err2023-09-01
err0
PREAI
errLe Chen; Pei-Hung Lin; Tristan Vanderbruggen; Chunhua Liao; Murali Emani; Bronis de Supinski
errShare
errSave
Faster sorting algorithms discovered using deep reinforcement learning
err2023-06-07
err0
errOAAI
errDaniel J. Mankowitz; Andrea Michi; Anton Zhernov; Marco Gelmi; Marco Selvi; Cosmin Paduraru; Edouard Leurent; Shariq Iqbal; Jean-Baptiste Lespiau; Alex Ahern; Thomas Köppe; Kevin Millikin; Stephen Gaffney; Sophie Elster; Jackson Broshear; Chris Gamble; Kieran Milan; Robert Tung; Minjae Hwang; Taylan Cemgil; Mohammadamin Barekatain; Yujia Li; Amol Mandhane; Thomas Hubert; Julian Schrittwieser; Demis Hassabis; Pushmeet Kohli; Martin Riedmiller; Oriol Vinyals; David Silver
errShare
errSave
A Comprehensive Overview of Large Language Models
err2025-08-19
err0
errOAAI
errHumza Naveed; Asad Ullah Khan; Shi Qiu; Muhammad Saqib; Saeed Anwar; Muhammad Usman; Naveed Akhtar; Nick Barnes; Ajmal Mian
errShare
errSave
errShare
errSave
OMPGPT: A Generative Pre-trained Transformer Model for OpenMP
err2024-08-26
err0
PREAI
errLe Chen; Arijit Bhattacharjee; Nesreen Ahmed; Niranjan Hasabnis; Gal Oren; Vy Vo; Ali Jannesari
errShare
errSave
researcher View more