arrow
Return

RunBugRun: An executable dataset for automated program repair

delete2026-03-24
delete0
PRE
AI
J
Julian Aron Prenner *
R
Romain Robbes
DOI:10.1007/s10664-025-10790-3delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recently, we can notice a transition to data-driven techniques in Automated Program Repair (APR), in particular towards deep neural networks. This entails training on hundreds of thousands or even millions of non-executable code fragments. We would like to bring more attention to an aspect of code often neglected in Neural Program Repair (NPR), namely its execution. Code execution has several significant advantages. It allows for test-based evaluation of candidate fixes and can provide valuable information to aid repair. In this work we present a fully executable dataset of 716,900 small buggy/fixed program pairs originally submitted to programming competition websites written in nine different programming languages. Along with the dataset we provide infrastructure to compile, safely execute and test programs as well as fine-grained bug-type labels. To give a point of reference, we provide basic evaluation results for CodeT5, Qwen3-Coder and Devstral baselines. With this dataset we follow several goals: we want to lift Neural Program Repair beyond fully static code representations, foster the use of execution-based features and, by including several different languages, counterbalance the predominance of Java in the current landscape of APR datasets and benchmarks.
Keywords:
Automated program repair
Neural program repair

Journal

Empirical Software Engineering cover
Empirical Software Engineering
IF:
3.6
Papers:
1.9K
Citations:
5.3K

Organization

F
Free University of Bozen-Bolzano
Scholars:
2.7K
Papers: 2.6K
Citations: 6
U
university of bordeaux
Scholars:
355
Papers: 144
Citations: 0