arrow
Return

Preprocessing noisy imbalanced datasets using SMOTE enhanced with fuzzy rough prototype selection

delete2014-09-01
delete69
PRE
AI
N
Nele Verbiest *
E
Enislay Ramentol
C
Chris Cornelis
F
Francisco Herrera
DOI:10.1016/j.asoc.2014.05.023delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The Synthetic Minority Over Sampling TEchnique (SMOTE) is a widely used technique to balance imbalanced data. In this paper we focus on improving SMOTE in the presence of class noise. Many improvements of SMOTE have been proposed, mostly cleaning or improving the data after applying SMOTE. Our approach differs from these approaches by the fact that it cleans the data before applying SMOTE, such that the quality of the generated instances is better. After applying SMOTE we also carry out data cleaning, such that instances (original or introduced by SMOTE) that badly fit in the new dataset are also removed. To this goal we propose two prototype selection techniques both based on fuzzy rough set theory. The first fuzzy rough prototype selection algorithm removes noisy instances from the imbalanced dataset, the second cleans the data generated by SMOTE. An experimental evaluation shows that our method improves existing preprocessing methods for imbalanced classification, especially in the presence of noise. (C) 2014 Elsevier B.V. All rights reserved.
Keywords:
Imbalanced classification
SMOTE
Prototype selection
Fuzzy rough set theory
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Applied Soft Computing cover
Applied Soft Computing
IF:
6.6
Papers:
1.4W
Citations:
4.8W

Organization

G
Ghent University
Scholars:
5.2W
Papers: 4.5W
Citations: 5.5W
K
King Abdulaziz University
Scholars:
2.0W
Papers: 1.9W
Citations: 3.3W
U
University of Granada
Scholars:
2.3W
Papers: 1.9W
Citations: 24
researcher View more organizations