arrow
返回

Using Permutation-Based Feature Importance for Improved Machine Learning Model Performance at Reduced Costs

delete2025-01-01
delete0
delete
OA
AI
M
M. Adam Khan
A
Asad Ali
J
Jahangir Khan
F
Fasee Ullah
M
Muhammad Faheem *
DOI:10.1109/ACCESS.2025.3544625delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In Software Quality Assurance (SQA), predicting defect-prone software modules is essential for ensuring software reliability and consistency. This task is commonly achieved through Machine Learning (ML) techniques, but improving model performance typically incurs significant computational costs. These high computational costs and uncertain payoffs make most Software engineering researchers reluctant to optimize ML models. This creates a need for novel techniques that can achieve near-optimal performance of hyperparameter settings while maintaining the computational efficiency of default settings. To address this, we employed five ML models, Decision Tree, Ranger, Random Forest, Support Vector Machine, and k-nearest Neighbors, and optimized their parameters using the random search technique. Our experiments covered six diverse Software Fault Prediction (SFP) datasets, encompassing various software features, application domains, and defect patterns, to evaluate the approach's generalizability and effectiveness. Moreover, the Permutation Feature Importance (PFI)-based model-agnostic method was employed to identify the top ten features most critical for model accuracy and efficiency. These selected features were used to retrain the ML models without hyperparameters (default settings) to determine whether similar performance could be achieved at low computational cost. The results show an average accuracy improvement of 77.39% and a 92.02% reduction in computational cost. The most important case attained a 99.25% accuracy improvement and a 96.77% cost reduction. Such results clearly show that PFI-based feature selection is capable of high performance at a fraction of computational cost, offering an efficient solution for software engineers to optimize ML models.
Keyword:
Computational modeling
Feature extraction
Accuracy
Computational efficiency
Predictive models
Optimization
Support vector machines
Random forests
Radio frequency
Decision trees
Model-agnostic techniques
permutation feature importance (PFI)
software fault prediction (SFP)
predictive accuracy
machine learning (ML)
computational cost
default settings
hyperparameter

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

V
vtt technical research center finland
学者数:
4.4K
论文数: 3.9K
被引数: 4
U
Universiti Teknologi Petronas
学者数:
5.4K
论文数: 4.6K
被引数: 5.9K
引用论文

引用论文

The Impact of Feature Importance Methods on the Interpretation of Defect Classifiers
err2022-07-01
err62
errOAAI
errRajbahadur, Gopi Krishnan; Wang, Shaowei; Oliva, Gustavo A.; Kamei, Yasutaka; Hassan, Ahmed E.
err分享
err收藏
err分享
err收藏
Feasibility of the Digital Retinography System Camera in the Pediatric Emergency Department
err2018-07-01
err0
PREAI
errYaron Ivan; Sriram Ramgopal; Margarita Cardenas-Villa; Daniel G. Winger; Li Wang; Melissa A. Vitale; Richard A. Saladino
err分享
err收藏
Multi-omics in Crohn's disease: New insights from inside
err2023-01-01
err0
errOAAI
errChenlu Mu; Qianjing Zhao; Qing Zhao; Lijiao Yang; Xiaoqi Pang; Tianyu Liu; Xiaomeng Li; Bangmao Wang; Shan-Yu Fung; Hailong Cao
err分享
err收藏
学者 查看更多内容