arrow
返回

A new sampling method for classifying imbalanced data based on support vector machine ensemble

delete2016-06-01
delete111
PRE
AI
C
Chuanxia Jian
高
高健 (Jian Gao)
Y
Yinhui Ao *
DOI:10.1016/j.neucom.2016.02.006delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The insufficient information from the minority examples cannot exactly represent the inherent structure of the dataset, which leads to a low prediction accuracy of the minority through the existing classification methods. The over- and under-sampling methods help to increase the prediction accuracy of the minority. However, the two methods either lose important information or add trivial information for classification, so as to affect the prediction accuracy of the minority. Therefore, a new different contribution sampling method (DCS) based on the contributions of the support vectors (SVs) and the nonsupport vectors (NSVs) to classification is proposed in this paper. The proposed DCS method applies different sampling methods for the SVs and the NSVs and uses the biased support vector machine (B-SVM) method to identify the SVs and the NSVs of an imbalanced data. Moreover, the synthetic minority over sampling technique (SMOTE) and the random under-sampling technique (RUS) are used in the proposed method to re-sample the SVs in the minority and the NSVs in the majority, respectively. Examples are labeled by the ensemble of support vector machine (SVMen). Experiments are carried out on the imbalanced dataset which is selected from UCI, AVU06a, Statlog, DP01a, JP98a and CWH03a repositories. Experimental results show that for the imbalanced datasets, the proposed DCS method achieves a better performance in the aspects of Receiver Operating Characteristic (ROC) curve than other methods. The proposed DCS method improves 20.80%, 5.97%, 8.66% and 9.35% in terms of the geometric mean prediction accuracy G(mean) as compared with that achieved by using the NS, the US, the SMOTE and the ROS, respectively. (C) 2016 Elsevier B.V. All rights reserved.
Keyword:
Imbalanced data
Sampling
Support vector machine
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Neurocomputing 封面图
Neurocomputing
IF:
6.5
论文数:
2.5W
被引数:
6.5W

机构

G
guangdong university of technology
学者数:
3.0W
论文数: 2.0W
被引数: 36
引用论文

引用论文

err分享
err收藏
Combustor Miniaturization with Liquid-Fuel Filming
err2003-11-11
err0
PREAI
errSimone Stanchi; Derek Dunn-Rankin; William Sirignano
err分享
err收藏
Urban Traffic Analysis through an UAV
err2014-02-01
err0
errOAAI
errGiuseppe Salvo; Luigi Caruso; Alessandro Scordo
err分享
err收藏
The strength of weak learnability
err1990-06-01
err0
errOAAI
errRobert E. Schapire
err分享
err收藏
err分享
err收藏
Analysing the classification of imbalanced data-sets with multiple classes: Binarization techniques and ad-hoc approaches
err2013-04-01
err301
PREAI
errFernandez, Alberto; Lopez, Victoria; Galar, Mikel; Jose del Jesus, Maria; Herrera, Francisco
err分享
err收藏
err分享
err收藏
Arsenic removal from aqueous solutions by adsorption using novel MIL-53(Fe) as a highly efficient adsorbent使用新型MIL-53(Fe) 作为高效吸附剂通过吸附从水溶液中去除砷
err2015-01-01
err0
PREAI
errTuan. A. Vu; Giang. H. Le; Canh. D. Dao; Lan. Q. Dang; Kien. T. Nguyen; Quang. K. Nguyen; Phuong. T. Dang; Hoa. T. K. Tran; Quang. T. Duong; Tuyen. V. Nguyen; Gun. D. Lee
err分享
err收藏
学者 查看更多内容