arrow
返回

Evaluating classifier performance with highly imbalanced Big Data

delete2023-04-11
delete22
delete
OA
AI
J
John Hancock *
T
Taghi M. Khoshgoftaar
J
Justin Johnson
DOI:10.1186/s40537-023-00724-5delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Using the wrong metrics to gauge classification of highly imbalanced Big Data may hide important information in experimental results. However, we find that analysis of metrics for performance evaluation and what they can hide or reveal is rarely covered in related works. Therefore, we address that gap by analyzing multiple popular performance metrics on three Big Data classification tasks. To the best of our knowledge, we are the first to utilize three new Medicare insurance claims datasets which became publicly available in 2021. These datasets are all highly imbalanced. Furthermore, the datasets are comprised of completely different data. We evaluate the performance of five ensemble learners in the Machine Learning task of Medicare fraud detection. Random Undersampling (RUS) is applied to induce five class ratios. The classifiers are evaluated with both the Area Under the Receiver Operating Characteristic Curve (AUC), and Area Under the Precision Recall Curve (AUPRC) metrics. We show that AUPRC provides a better insight into classification performance. Our findings reveal that the AUC metric hides the performance impact of RUS. However, classification results in terms of AUPRC show RUS has a detrimental effect. We show that, for highly imbalanced Big Data, the AUC metric fails to capture information about precision scores and false positive counts that the AUPRC metric reveals. Our contribution is to show AUPRC is a more effective metric for evaluating the performance of classifiers when working with highly imbalanced Big Data.
Keyword:
Extremely randomized trees
XGBoost
Class imbalance
Big Data
Undersampling
AUC
AUPRC

期刊

Journal of Big Data 封面图
Journal of Big Data
IF:
6.4
论文数:
1.5K
被引数:
1.1W

机构

State University System of Florida 封面图
State University System of Florida
学者数:
12.7W
论文数: 10.9W
被引数: 130
引用论文

引用论文

err分享
err收藏
Detecting web attacks using random undersampling and ensemble learners
err2021-05-27
err35
errOAAI
errZuech, Richard; Hancock, John; Khoshgoftaar, Taghi M.
err分享
err收藏
Bagging predictorsBagging预测器
err1996-08-01
err1.0W
PREAI
errBreiman, L
err分享
err收藏
Severely imbalanced Big Data challenges: investigating data sampling approaches
err2019-11-30
err73
errOAAI
errHasanin, Tawfiq; Khoshgoftaar, Taghi M.; Leevy, Joffrey L.; Bauder, Richard A.
err分享
err收藏
Medicare fraud detection using neural networks
err2019-07-18
err186
errOAAI
errJohnson, Justin M.; Khoshgoftaar, Taghi M.
err分享
err收藏
Pyroelectric properties of Mn-doped 94.6Na0.5Bi0.5TiO3-5.4BaTiO3 lead-free single crystals
err2014-02-21
err0
PREAI
errRenbing Sun; Jinzhi Wang; Fang Wang; Tangfu Feng; Yanlong Li; Zhenhua Chi; Xiangyong Zhao; Haosu Luo
err分享
err收藏
Apache Spark: A Unified Engine for Big Data ProcessingApache Spark: 用于大数据处理的统一引擎
err2016-10-28
err1.7K
PREAI
errZaharia, Matei; Xin, Reynold S.; Wendell, Patrick; Das, Tathagata; Armbrust, Michael; Dave, Ankur; Meng, Xiangrui; Rosen, Josh; Venkataraman, Shivaram; Franklin, Michael J.; Ghodsi, Ali; Gonzalez, Joseph; Shenker, Scott; Stoica, Ion
err分享
err收藏
err分享
err收藏
学者 查看更多内容