arrow
Return

A new sampling method for classifying imbalanced data based on support vector machine ensemble

delete2016-06-01
delete111
PRE
AI
C
Chuanxia Jian
高
高健 (Jian Gao)
Y
Yinhui Ao *
DOI:10.1016/j.neucom.2016.02.006delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The insufficient information from the minority examples cannot exactly represent the inherent structure of the dataset, which leads to a low prediction accuracy of the minority through the existing classification methods. The over- and under-sampling methods help to increase the prediction accuracy of the minority. However, the two methods either lose important information or add trivial information for classification, so as to affect the prediction accuracy of the minority. Therefore, a new different contribution sampling method (DCS) based on the contributions of the support vectors (SVs) and the nonsupport vectors (NSVs) to classification is proposed in this paper. The proposed DCS method applies different sampling methods for the SVs and the NSVs and uses the biased support vector machine (B-SVM) method to identify the SVs and the NSVs of an imbalanced data. Moreover, the synthetic minority over sampling technique (SMOTE) and the random under-sampling technique (RUS) are used in the proposed method to re-sample the SVs in the minority and the NSVs in the majority, respectively. Examples are labeled by the ensemble of support vector machine (SVMen). Experiments are carried out on the imbalanced dataset which is selected from UCI, AVU06a, Statlog, DP01a, JP98a and CWH03a repositories. Experimental results show that for the imbalanced datasets, the proposed DCS method achieves a better performance in the aspects of Receiver Operating Characteristic (ROC) curve than other methods. The proposed DCS method improves 20.80%, 5.97%, 8.66% and 9.35% in terms of the geometric mean prediction accuracy G(mean) as compared with that achieved by using the NS, the US, the SMOTE and the ROS, respectively. (C) 2016 Elsevier B.V. All rights reserved.
Keywords:
Imbalanced data
Sampling
Support vector machine
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

G
guangdong university of technology
Scholars:
3.0W
Papers: 2.0W
Citations: 36
Cited Papers

Cited Papers

errShare
errSave
Combustor Miniaturization with Liquid-Fuel Filming
err2003-11-11
err0
PREAI
errSimone Stanchi; Derek Dunn-Rankin; William Sirignano
errShare
errSave
Urban Traffic Analysis through an UAV
err2014-02-01
err0
errOAAI
errGiuseppe Salvo; Luigi Caruso; Alessandro Scordo
errShare
errSave
The strength of weak learnability
err1990-06-01
err0
errOAAI
errRobert E. Schapire
errShare
errSave
errShare
errSave
Analysing the classification of imbalanced data-sets with multiple classes: Binarization techniques and ad-hoc approaches
err2013-04-01
err301
PREAI
errFernandez, Alberto; Lopez, Victoria; Galar, Mikel; Jose del Jesus, Maria; Herrera, Francisco
errShare
errSave
errShare
errSave
Arsenic removal from aqueous solutions by adsorption using novel MIL-53(Fe) as a highly efficient adsorbent
err2015-01-01
err0
PREAI
errTuan. A. Vu; Giang. H. Le; Canh. D. Dao; Lan. Q. Dang; Kien. T. Nguyen; Quang. K. Nguyen; Phuong. T. Dang; Hoa. T. K. Tran; Quang. T. Duong; Tuyen. V. Nguyen; Gun. D. Lee
errShare
errSave
researcher View more