arrow
Return

Label correlation guided borderline oversampling for imbalanced multi-label data learning

delete2023-11-01
delete7
PRE
AI
K
Kai Zhang
P
Peng Cao *
W
Wei Liang
J
Jinzhu Yang
W
Weiping Li
O
Osmar R. Zaı̈ane
DOI:10.1016/j.knosys.2023.110938delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Multi-label data classification has received much attention due to its wide range of application domains. Unfortunately, a class imbalance problem often occurs in multi-label datasets, causing challenges for classification algorithms. Oversampling is one of the most important approaches, as it generates minority label instances to balance the class distribution. However, existing oversampling methods ignore existing label correlations, resulting in the generation of inappropriate synthetic minority samples and making multi-label data classification tasks harder. In this work, we propose an oversampling method that considers label correlations and identifies two critical boundary regions for generating synthetic minority samples. Moreover, we propose a weighting strategy to assign weights to these instances based on their distance information. To evaluate the performance of our proposed method, we conducted experiments on sixteen public datasets. The results show that our approach outperforms the state-of-the-art approaches in terms of various assessment metrics, such as Macro F1 and Macro AUC. The code is available at https://github.com/IntelliDAL/Multi-label/tree/main/LCOS. (c) 2023 Elsevier B.V. All rights reserved.
Keywords:
Class imbalance
Multi-label data classification
Oversampling
Label correlation
Critical boundary regions

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

U
university of alberta
Scholars:
5.1W
Papers: 4.9W
Citations: 65
N
northeastern university - china
Scholars:
3.1W
Papers: 2.7W
Citations: 37