arrow
返回

Class-imbalanced classifiers for high-dimensional data

delete2012-03-09
delete214
PRE
AI
W
Wei‐Jiun Lin
J
James J. Chen *
DOI:10.1093/bib/bbs006delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
A class-imbalanced classifier is a decision rule to predict the class membership of new samples from an available data set where the class sizes differ considerably. When the class sizes are very different, most standard classification algorithms may favor the larger (majority) class resulting in poor accuracy in the minority class prediction. A class-imbalanced classifier typically modifies a standard classifier by a correction strategy or by incorporating a new strategy in the training phase to account for differential class sizes. This article reviews and evaluates some most important methods for class prediction of high-dimensional imbalanced data. The evaluation addresses the fundamental issues of the class-imbalanced classification problem: imbalance ratio, small disjuncts and overlap complexity, lack of data and feature selection. Four class-imbalanced classifiers are considered. The four classifiers include three standard classification algorithms each coupled with an ensemble correction strategy and one support vector machines (SVM)-based correction classifier. The three algorithms are (i) diagonal linear discriminant analysis (DLDA), (ii) random forests (RFs) and (ii) SVMs. The SVM-based correction classifier is SVM threshold adjustment (SVM-THR). A Monte-Carlo simulation and five genomic data sets were used to illustrate the analysis and address the issues. The SVM-ensemble classifier appears to perform the best when the class imbalance is not too severe. The SVM-THR performs well if the imbalance is severe and predictors are highly correlated. The DLDA with a feature selection can perform well without using the ensemble correction.
Keyword:
class-imbalanced prediction
feature selection
lack of data
performance metrics
threshold adjustment
under-sampling ensemble
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Briefings in Bioinformatics 封面图
Briefings in Bioinformatics
IF:
7.7
论文数:
5.8K
被引数:
2.7W

机构

U
us food & drug administration (fda)
学者数:
1.4W
论文数: 9.3K
被引数: 3
F
Feng Chia University
学者数:
3.5K
论文数: 3.7K
被引数: 2.6K
引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
err分享
err收藏
Diffuse large B-cell lymphoma outcome prediction by gene-expression profiling and supervised machine learning
err2002-01-01
err2.1K
PREAI
errShipp, MA; Ross, KN; Tamayo, P; Weng, AP; Kutok, JL; Aguiar, RCT; Gaasenbeek, M; Angelo, M; Reich, M; Pinkus, GS; Ray, TS; Koval, MA; Last, KW; Norton, A; Lister, TA; Mesirov, J; Neuberg, DS; Lander, ES; Aster, JC; Golub, TR
err分享
err收藏
Comparison of the effects of isradipine and lisinopril on left ventricular structure and function in essential hypertension
err1992-05-01
err0
PREAI
errEdith C. Bielen; Robert H. Fagard; Paul J. Lijnen; Tikma B. Tjandra-Maga; René Verbessert; Antoon K. Amery
err分享
err收藏
Cancer Statistics, 2010癌症统计,2010
err2010-07-07
err1.2W
errOAAI
errJemal, Ahmedin; Siegel, Rebecca; Xu, Jiaquan; Ward, Elizabeth
err分享
err收藏
Oligonucleotide microarray for prediction of early intrahepatic recurrence of hepatocellular carcinoma after curative resection寡核苷酸芯片预测肝癌切除术后早期肝内复发
errLANCET
IF88.5
err2003-03-01
err472
PREAI
errIizuka, N; Oka, M; Yamada-Okabe, H; Nishida, M; Maeda, Y; Mori, N; Takao, T; Tamesa, T; Tangoku, A; Tabuchi, H; Hamada, K; Nakayama, H; Ishitsuka, H; Miyamoto, T; Hirabayashi, A; Uchimura, S; Hamamoto, Y
err分享
err收藏
学者 查看更多内容