arrow
Return

Class-imbalanced classifiers for high-dimensional data

delete2012-03-09
delete214
PRE
AI
W
Wei‐Jiun Lin
J
James J. Chen *
DOI:10.1093/bib/bbs006delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
A class-imbalanced classifier is a decision rule to predict the class membership of new samples from an available data set where the class sizes differ considerably. When the class sizes are very different, most standard classification algorithms may favor the larger (majority) class resulting in poor accuracy in the minority class prediction. A class-imbalanced classifier typically modifies a standard classifier by a correction strategy or by incorporating a new strategy in the training phase to account for differential class sizes. This article reviews and evaluates some most important methods for class prediction of high-dimensional imbalanced data. The evaluation addresses the fundamental issues of the class-imbalanced classification problem: imbalance ratio, small disjuncts and overlap complexity, lack of data and feature selection. Four class-imbalanced classifiers are considered. The four classifiers include three standard classification algorithms each coupled with an ensemble correction strategy and one support vector machines (SVM)-based correction classifier. The three algorithms are (i) diagonal linear discriminant analysis (DLDA), (ii) random forests (RFs) and (ii) SVMs. The SVM-based correction classifier is SVM threshold adjustment (SVM-THR). A Monte-Carlo simulation and five genomic data sets were used to illustrate the analysis and address the issues. The SVM-ensemble classifier appears to perform the best when the class imbalance is not too severe. The SVM-THR performs well if the imbalance is severe and predictors are highly correlated. The DLDA with a feature selection can perform well without using the ensemble correction.
Keywords:
class-imbalanced prediction
feature selection
lack of data
performance metrics
threshold adjustment
under-sampling ensemble
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Briefings in Bioinformatics cover
Briefings in Bioinformatics
IF:
7.7
Papers:
5.8K
Citations:
2.7W

Organization

U
us food & drug administration (fda)
Scholars:
1.4W
Papers: 9.3K
Citations: 3
F
Feng Chia University
Scholars:
3.5K
Papers: 3.7K
Citations: 2.6K
Cited Papers

Cited Papers

err
IF0
err
err0
PREAI
err
errShare
errSave
errShare
errSave
Diffuse large B-cell lymphoma outcome prediction by gene-expression profiling and supervised machine learning
err2002-01-01
err2.1K
PREAI
errShipp, MA; Ross, KN; Tamayo, P; Weng, AP; Kutok, JL; Aguiar, RCT; Gaasenbeek, M; Angelo, M; Reich, M; Pinkus, GS; Ray, TS; Koval, MA; Last, KW; Norton, A; Lister, TA; Mesirov, J; Neuberg, DS; Lander, ES; Aster, JC; Golub, TR
errShare
errSave
Comparison of the effects of isradipine and lisinopril on left ventricular structure and function in essential hypertension
err1992-05-01
err0
PREAI
errEdith C. Bielen; Robert H. Fagard; Paul J. Lijnen; Tikma B. Tjandra-Maga; René Verbessert; Antoon K. Amery
errShare
errSave
Cancer Statistics, 2010
err2010-07-07
err1.2W
errOAAI
errJemal, Ahmedin; Siegel, Rebecca; Xu, Jiaquan; Ward, Elizabeth
errShare
errSave
Oligonucleotide microarray for prediction of early intrahepatic recurrence of hepatocellular carcinoma after curative resection
errLANCET
IF88.5
err2003-03-01
err472
PREAI
errIizuka, N; Oka, M; Yamada-Okabe, H; Nishida, M; Maeda, Y; Mori, N; Takao, T; Tamesa, T; Tangoku, A; Tabuchi, H; Hamada, K; Nakayama, H; Ishitsuka, H; Miyamoto, T; Hirabayashi, A; Uchimura, S; Hamamoto, Y
errShare
errSave
researcher View more