arrow
Return

How repeated data points affect bug prediction performance: A case study

delete2016-12-01
delete4
PRE
AI
M
Muhammed Maruf Öztürk *
A
Ahmet Zengi̇n
DOI:10.1016/j.asoc.2016.08.002delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In defect prediction studies, open-source and real-world defect data sets are frequently used. The quality of these data sets is one of the main factors affecting the validity of defect prediction methods. One of the issues is repeated data points in defect prediction data sets. The main goal of the paper is to explore how low-level metrics are derived. This paper also presents a cleansing algorithm that removes repeated data points from defect data sets. The method was applied on 20 data sets, including five open source sets, and area under the curve (AUC) and precision performance parameters have been improved by 4.05% and 6.7%, respectively. In addition, this work discusses how static code metrics should be used in bug prediction. The study provides tips to obtain better defect prediction results. (C) 2016 Elsevier B.V. All rights reserved.
Keywords:
Bug prediction
Repeated data
Software metricsa
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Applied Soft Computing cover
Applied Soft Computing
IF:
6.6
Papers:
1.4W
Citations:
4.8W

Organization

No organization information available
Cited Papers

Cited Papers

On the usefulness of ownership metrics in open-source software projects
err2015-08-01
err16
errOAAI
errFoucault, Matthieu; Teyton, Cedric; Lo, David; Blanc, Xavier; Falleri, Jean-Remy
errShare
errSave
Data Quality: Some Comments on the NASA Software Defect Datasets
err2013-09-01
err375
errOAAI
errShepperd, Martin; Song, Qinbao; Sun, Zhongbin; Mair, Carolyn
errShare
errSave
errShare
errSave
errShare
errSave
errShare
errSave
researcher View more