arrow
返回

An interpretable automated feature engineering framework for improving logistic regression

delete2024-03-01
delete2
PRE
AI
M
Mucan Liu
郭
郭崇慧 (Chonghui Guo) *
DOI:10.1016/j.asoc.2024.111269delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Although black -box models such as ensemble learning models often provide better predictive performance than intrinsic interpretable models such as logistic regression, black -box models are not still applicable due to the lack of interpretability. Recently, there has been an explosion of work on explainable machine learning techniques, which utilize external algorithms or models to explain the behavior of black -box models. However, it is problematic to explain the black -box model behavior because the explanation provided might not reveal the real mechanism or decision process of black -box models. In this study, instead of using explainable machine learning techniques, an automated feature engineering task was formulated to help logistic regression achieve predictive performance comparable to or even better than black -box models while maintaining interpretability. In this paper, an INterpretable Automated Feature ENgineering (INAFEN) framework was designed for logistic regression. This framework automatically transforms the nonlinear relationships between numerical features and labels into linear relationships, conducts feature cross through association rule mining, and distills knowledge from black -box models. A case study was performed on gastric survival prediction to present the rationality of the feature transformations through INAFEN and benchmark experiments to show the validity of INAFEN. Experimental results on 10 classification tasks demonstrated that INAFEN achieved an average ranking of 2.60 in area under the ROC curve (AUROC), 3.35 in area under the PR curve (AUROC), 3.70 in F1 score and 3.00 in Brier score (among 13 models), outperforming other interpretable baselines and even black -box models. In addition, the interpretability measurement of INAFEN is significantly better than that of black -box models.
Keyword:
Interpretable machine learning
Feature engineering
Automated machine learning
Knowledge distillation

期刊

Applied Soft Computing 封面图
Applied Soft Computing
IF:
6.6
论文数:
1.4W
被引数:
4.8W

机构

D
Dalian University of Technology
学者数:
6.0W
论文数: 4.4W
被引数: 5.5W
引用论文

引用论文

err分享
err收藏
Interpretable machine learning: Fundamental principles and 10 grand challenges可解释的机器学习: 基本原则和十大挑战
err2022-01-01
err290
errOAAI
errRudin, Cynthia; Chen, Chaofan; Chen, Zhi; Huang, Haiyang; Semenova, Lesia; Zhong, Chudi
err分享
err收藏
Using the Linearized Navier-Stokes Equations to Model Acoustic Liners
err2018-06-24
err0
PREAI
errMads Jakob Herring Jensen; Elin Svensson; Kirill Shaposhnikov
err分享
err收藏
A review of uncertainty quantification in deep learning: Techniques, applications and challenges深度学习中的不确定性量化: 技术、应用与挑战
err2021-12-01
err1.2K
errOAAI
errAbdar, Moloud; Pourpanah, Farhad; Hussain, Sadiq; Rezazadegan, Dana; Liu, Li; Ghavamzadeh, Mohammad; Fieguth, Paul; Cao, Xiaochun; Khosravi, Abbas; Acharya, U. Rajendra; Makarenkov, Vladimir; Nahavandi, Saeid
err分享
err收藏
A multi-level classification and modified PSO clustering based ensemble approach for credit scoring
err2021-11-01
err8
PREAI
errSingh, Indu; Kumar, Narendra; Srinivasa, K. G.; Maini, Shivam; Ahuja, Umang; Jain, Siddhant
err分享
err收藏
学者 查看更多内容