arrow
返回

A novel progressively undersampling method based on the density peaks sequence for imbalanced data

delete2021-02-01
delete53
PRE
AI
刘
刘华文 (Huawen Liu) *
S
Shouzhen Zeng *
W
Wen Li
DOI:10.1016/j.knosys.2020.106689delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Undersampling is a widely used resampling technique for imbalanced data. As traditional undersampling techniques, typically making majority and minority classes in imbalanced data into the same scale, tend to miss valuable information, many strategies like clustering have been developed. However, two essential problems still remain and require more efforts to be put; that is, which and how many instances should be extracted in undersampling. To alleviate these two problems, in this paper we propose a novel undersampling method for imbalanced data. It exploits a sequence of density peaks to progressively extract instances from the majority classes of the imbalanced data. Specifically, two factors are introduced to measure the importance degree of each instance in the majority classes. With these two factors, we generate a sampling sequence based on the importance of instances for classification. Furthermore, the optimal undersampling size of the majority classes is automatically determined by progressively extracting the important instances from the sequence. To evaluate the effectiveness of the proposed method, a series of experiments comparing to six popular undersampling methods were conducted on 40 public benchmark datasets. The experimental results show that the performance of the proposed undersampling method is superior to the state-of-the-art undersampling methods. (C) 2020 Elsevier B.V. All rights reserved.
Keyword:
Progressive undersampling
Density peaks sequence
Importance degree
Optimal undersampling size
Imbalanced data
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

K
Knowledge-Based Systems
IF:
7.6
论文数:
1.3W
被引数:
4.5W

机构

Z
Zhejiang Normal University
学者数:
1.3W
论文数: 8.4K
被引数: 1.2W
C
Curtin University
学者数:
1.5W
论文数: 1.8W
被引数: 2.8W
N
Ningbo University
学者数:
2.6W
论文数: 1.8W
被引数: 2.4W
学者 查看更多机构
引用论文

引用论文

Indinavir and Rifabutin Drug Interactions in Healthy Volunteers
err2013-03-08
err0
PREAI
errWalter K. Kraft; Jacqueline B. McCrea; Gregory A. Winchell; Alexandra Carides; Richard Lowry; Eric J. Woolf; Sandra E. Kusma; Paul J. Deutsch; Howard E. Greenberg; Scott A. Waldman
err分享
err收藏
Diversified Sensitivity-Based Undersampling for Imbalance Classification Problems
err2015-11-01
err181
PREAI
errNg, Wing W. Y.; Hu, Junjie; Yeung, Daniel S.; Yin, Shaohua; Roli, Fabio
err分享
err收藏
A comprehensive data level analysis for cancer diagnosis on imbalanced data
err2019-02-01
err194
PREAI
errFotouhi, Sara; Asadi, Shahrokh; Kattan, Michael W.
err分享
err收藏
err分享
err收藏
Analysing the classification of imbalanced data-sets with multiple classes: Binarization techniques and ad-hoc approaches
err2013-04-01
err301
PREAI
errFernandez, Alberto; Lopez, Victoria; Galar, Mikel; Jose del Jesus, Maria; Herrera, Francisco
err分享
err收藏
Cost-sensitive KNN classification
err2020-05-01
err125
PREAI
errZhang, Shichao
err分享
err收藏
Response of Short Food Supply Chains in Western Balkan Countries to the COVID Crisis: A Case Study in the Honey Sector
err2024-04-02
err0
errOAAI
errVesna Paraušić; Etleva Muça Dashi; Jonel Subić; Iwona Pomianek; Bojana Bekić Šarić
err分享
err收藏
Multi-Imbalance: An open-source software for multi-class imbalance learning多不平衡: 用于多类不平衡学习的开源软件
err2019-06-01
err136
PREAI
errZhang, Chongsheng; Bi, Jingjun; Xu, Shixin; Ramentol, Enislay; Fan, Gaojuan; Qiao, Baojun; Fujita, Hamido
err分享
err收藏
学者 查看更多内容