arrow
Return

Improving Ranking-Oriented Defect Prediction Using a Cost-Sensitive Ranking SVM

delete2020-03-01
delete44
PRE
AI
X
Xiao Yu
J
Jin Liu *
J
Jacky Keung *
Q
Qing Li
K
Kwabena Ebo Bennin
周旭 cover
周旭 (Zhou Xu)
J
Jun Ping Wang
崔晓晖 (Xiaohui Cui)
DOI:10.1109/TR.2019.2931559delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Context: Ranking-oriented defect prediction (RODP) ranks software modules to allocate limited testing resources to each module according to the predicted number of defects. Most RODP methods overlook that ranking a module with more defects incorrectly makes it difficult to successfully find all of the defects in the module due to fewer testing resources being allocated to the module, which results in much higher costs than incorrectly ranking the modules with fewer defects, and the numbers of defects in software modules are highly imbalanced in defective software datasets. Cost-sensitive learning is an effective technique in handling the cost issue and data imbalance problem for software defect prediction. However, the effectiveness of cost-sensitive learning has not been investigated in RODP models. Aims: In this article, we propose a cost-sensitive ranking support vector machine (SVM) (CSRankSVM) algorithm to improve the performance of RODP models. Method: CSRankSVM modifies the loss function of the ranking SVM algorithm by adding two penalty parameters to address both the cost issue and the data imbalance problem. Additionally, the loss function of the CSRankSVM is optimized using a genetic algorithm. Results: The experimental results for 11 project datasets with 41 releases show that CSRankSVM achieves 1.12%-15.68% higher average fault percentile average (FPA) values than the five existing RODP methods (i.e., decision tree regression, linear regression, Bayesian ridge regression, ranking SVM, and learning-to-rank (LTR)) and 1.08%-15.74% higher average FPA values than the four data imbalance learning methods (i.e., random undersampling and a synthetic minority oversampling technique; two data resampling methods; RankBoost, an ensemble learning method; IRSVM, a CSRankSVM method for information retrieval). Conclusion: CSRankSVM is capable of handling the cost issue and data imbalance problem in RODP methods and achieves better performance. Therefore, CSRankSVM is recommended as an effective method for RODP.
Keywords:
Support vector machines
Software
Prediction algorithms
Predictive models
Testing
Software algorithms
Computer science
Cost-sensitive learning
data imbalance
ranking-oriented defect prediction (RODP)
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Reliability cover
IEEE Transactions on Reliability
IF:
5.7
Papers:
2.7K
Citations:
8.5K

Organization

H
hong kong polytechnic university
Scholars:
3.0W
Papers: 4.1W
Citations: 921
I
institute of automation, cas
Scholars:
2.2K
Papers: 2.1K
Citations: 2
C
City University of Hong Kong
Scholars:
2.3W
Papers: 3.0W
Citations: 6.1W
B
blekinge institute technology
Scholars:
744
Papers: 760
Citations: 5
W
wuhan university
Scholars:
8.1W
Papers: 5.8W
Citations: 70
C
chinese academy of sciences
Scholars:
56.5W
Papers: 44.9W
Citations: 704
researcher View more organizations