arrow
返回

Classifying imbalanced data sets using similarity based hierarchical decomposition

delete2015-05-01
delete109
delete
OA
AI
C
Cigdem Beyan *
R
Robert B. Fisher
DOI:10.1016/j.patcog.2014.10.032delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Classification of data is difficult if the data is imbalanced and classes are overlapping. In recent years, more research has started to focus on classification of imbalanced data since real world data is often skewed. Traditional methods are more successful with classifying the class that has the most samples (majority class) compared to the other classes (minority classes). For the classification of imbalanced data sets, different methods are available, although each has some advantages and shortcomings. In this study, we propose a new hierarchical decomposition method for imbalanced data sets which is different from previously proposed solutions to the class imbalance problem. Additionally, it does not require any data pre-processing step as many other solutions need. The new method is based on clustering and outlier detection. The hierarchy is constructed using the similarity of labeled data subsets at each level of the hierarchy with different levels being built by different data and feature subsets. Clustering is used to partition the data while outlier detection is utilized to detect minority class samples. The comparison of the proposed method with state of art the methods using 20 public imbalanced data sets and 181 synthetic data sets showed that the proposed method's classification performance is better than the state of art methods. It is especially successful if the minority class is sparser than the majority class. It has accurate performance even when classes have sub-varieties and minority and majority classes are overlapping. Moreover, its performance is also good when the class imbalance ratio is low, i.e. classes are more imbalanced. (C) 2014 Elsevier Ltd. All rights reserved.
Keyword:
Class imbalance problem
Hierarchical decomposition
Clustering
Outlier detection
Minority-majority classes
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Pattern Recognition 封面图
Pattern Recognition
IF:
7.6
论文数:
1.3W
被引数:
4.5W

机构

U
University of Edinburgh
学者数:
5.2W
论文数: 4.6W
被引数: 71
引用论文

引用论文

Combustor Miniaturization with Liquid-Fuel Filming
err2003-11-11
err0
PREAI
errSimone Stanchi; Derek Dunn-Rankin; William Sirignano
err分享
err收藏
E-Government Innovation, Financial Disclosure, and Public Sector Accounts
err2022-11-17
err0
PREAI
errSuwastika Naidu; Fang Zhao; Anand Chand; Arvind Patel; Atishwar Pandaram
err分享
err收藏
Mine Classification With Imbalanced Data
err2009-07-01
err75
PREAI
errWilliams, David P.; Myers, Vincent; Silvious, Miranda Schatten
err分享
err收藏
Learning multi-label scene classification学习多标签场景分类
err2004-09-01
err2.0K
PREAI
errBoutell, MR; Luo, JB; Shen, XP; Brown, CM
err分享
err收藏
err分享
err收藏
Arsenic removal from aqueous solutions by adsorption using novel MIL-53(Fe) as a highly efficient adsorbent使用新型MIL-53(Fe) 作为高效吸附剂通过吸附从水溶液中去除砷
err2015-01-01
err0
PREAI
errTuan. A. Vu; Giang. H. Le; Canh. D. Dao; Lan. Q. Dang; Kien. T. Nguyen; Quang. K. Nguyen; Phuong. T. Dang; Hoa. T. K. Tran; Quang. T. Duong; Tuyen. V. Nguyen; Gun. D. Lee
err分享
err收藏
Neuronal-binding antibodies from patients with antiphospholipid syndrome induce cognitive deficits following intrathecal passive transfer
err2003-06-01
err0
PREAI
errY Shoenfeld; A Nahum; A D Korczyn; M Dano; R Rabinowitz; O Beilin; C G Pick; L Leider-Trejo; L Kalashnikova; M Blank; J Chapman
err分享
err收藏
学者 查看更多内容