arrow
返回

Improving transfer learning for software cross-project defect prediction

delete2024-04-24
delete3
delete
OA
AI
O
Osayande Pascal Omondiagbe *
S
Sherlock A. Licorish
S
Stephen G. MacDonell
DOI:10.1007/s10489-024-05459-1delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Software cross-project defect prediction (CPDP) makes use of cross-project (CP) data to overcome the lack of data necessary to train well-performing software defect prediction (SDP) classifiers in the early stage of new software projects. Since the CP data (known as the source) may be different from the new project's data (known as the target), this makes it difficult for CPDP classifiers to perform well. In particular, it is a mismatch of data distributions between source and target that creates this difficulty. Transfer learning-based CPDP classifiers are designed to minimize these distribution differences. The first Transfer learning-based CPDP classifiers treated these differences equally, thereby degrading prediction performance. To this end, recent research has the Weighted Balanced Distribution Adaptation (W-BDA) method to leverage the importance of both distribution differences to improve classification performance. Although W-BDA has been shown to improve model performance in CPDP and tackle the class imbalance by balancing the class proportion of each domain, research to date has failed to consider model performance in light of increasing target data. We provide the first investigation studying the effects of increasing the target data when leveraging the importance of both distribution differences. We extend the initial W-BDA method and call this extension the W-BDA + \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathbf {<^>{+}}$$\end{document} method. To evaluate the effectiveness of W-BDA + \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathbf {<^>{+}}$$\end{document} for improving CPDP performance, we conduct eight experiments on 18 projects from four datasets, where data sampling was performed with different sampling methods. Data sampling was only performed on the baseline methods and not on our proposed W-BDA + \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathbf {<^>{+}}$$\end{document} and the original W-BDA because data sampling issues do not exist for these two methods. We evaluate our method using four complementary indicators (i.e., Balanced Accuracy, AUC, F-measure and G-Measure). Our findings reveal an average improvement of 6%, 7.5%, 10% and 12% for these four indicators when W-BDA + \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathbf {<^>{+}}$$\end{document} is compared to the original W-BDA and five other baseline methods (for all four of the sampling methods used). Also, as the target to source ratio is increased with different sampling methods, we observe a decrease in performance for the original W-BDA, with our W-BDA + \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\mathbf {<^>{+}}$$\end{document} approach outperforming the original W-BDA in most cases. Our results highlight the importance of having an awareness of the effect of the increasing availability of target data in CPDP scenarios when using a method that can handle the class imbalance problem.
Keyword:
Transfer learning
Cross-project defect prediction
Weighted balance distribution

期刊

Applied Intelligence 封面图
Applied Intelligence
IF:
3.5
论文数:
7.6K
被引数:
1.7W

机构

A
Auckland University of Technology
学者数:
4.0K
论文数: 4.4K
被引数: 4.7K
U
university of otago
学者数:
1.8W
论文数: 1.6W
被引数: 15
引用论文

引用论文

Creep tests on notched specimens of copper
err2018-10-01
err0
PREAI
errFangfei Sui; Rolf Sandström; Rui Wu
err分享
err收藏
Crack Development in Cementitious Materials under Impact Loading
err2011-02-25
err0
PREAI
errSidney Mindess; Nemy P. Banthia; Andrew Ritter; Jan P. Skalny
err分享
err收藏
Rent-seeking under a weak institutional environment
err2009-09-01
err0
PREAI
errDavide Infante; Janna Smirnova
err分享
err收藏
Occupational injuries in California's health care and social assistance industry, 2009 to 2018
err2021-06-06
err0
errOAAI
errKerri Wizner; Fraser W. Gaspar; Adriane Biggio; Steve Wiesner
err分享
err收藏
Driving Current through Single Organic Molecules
err2002-04-10
err0
errOAAI
errJ. Reichert; R. Ochs; D. Beckmann; H. B. Weber; M. Mayor; H. v. Löhneysen
err分享
err收藏
学者 查看更多内容