arrow
返回

Bugs in machine learning-based systems: a faultload benchmark

delete2023-04-05
delete9
PRE
AI
M
Mohammad Mehdi Morovati *
A
Amin Nikanjam
F
Foutse Khomh
Z
Zhen Ming Jiang
DOI:10.1007/s10664-023-10291-1delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The rapid escalation of applying Machine Learning (ML) in various domains has led to paying more attention to the quality of ML components. There is then a growth of techniques and tools aiming at improving the quality of ML components and integrating them into the ML-based system safely. Although most of these tools use bugs' lifecycle, there is no standard benchmark of bugs to assess their performance, compare them and discuss their advantages and weaknesses. In this study, we firstly investigate the reproducibility and verifiability of the bugs in ML-based systems and show the most important factors in each one. Then, we explore the challenges of generating a benchmark of bugs in ML-based software systems and provide a bug benchmark namely defect4ML that satisfies all criteria of standard benchmark, i.e. relevance, reproducibility, fairness, verifiability, and usability. This faultload benchmark contains 100 bugs reported by ML developers in GitHub and Stack Overflow, using two of the most popular ML frameworks: TensorFlow and Keras. defect4ML also addresses important challenges in Software Reliability Engineering of ML-based software systems, like: 1) fast changes in frameworks, by providing various bugs for different versions of frameworks, 2) code portability, by delivering similar bugs in different ML frameworks, 3) bug reproducibility, by providing fully reproducible bugs with complete information about required dependencies and data, and 4) lack of detailed information on bugs, by presenting links to the bugs' origins. defect4ML can be of interest to ML-based systems practitioners and researchers to assess their testing tools and techniques.
Keyword:
Benchmark
Machine learning-based system
Software bug
Software reliability engineering
Software testing

期刊

Empirical Software Engineering 封面图
Empirical Software Engineering
IF:
3.6
论文数:
2.0K
被引数:
5.3K

机构

U
universite de montreal
学者数:
4.6W
论文数: 3.8W
被引数: 46
P
Polytechnique Montreal
学者数:
3.7K
论文数: 3.4K
被引数: 42
引用论文

引用论文

Neural representation of goal direction in the monarch butterfly brain
err
IF0
err2022-10-18
err0
errOAAI
errM. Jerome Beetz; Christian Kraus; Basil el Jundi
err分享
err收藏
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
学者 查看更多内容