arrow
Return

An imbalanced binary classification method based on contrastive learning using multi-label confidence comparisons within sample- neighbors pair

delete2023-01-01
delete8
PRE
AI
X
Xin Gao *
X
Xin Jia
J
Jing Liu
X
Xinping Diao
B
Bing Xue
Z
Zijian Huang
K
Kangsheng Li
DOI:10.1016/j.neucom.2022.10.069delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
For imbalanced classification, data-level methods can achieve inter-class balance, but the samples gener-ated do not contain new information and cannot avoid the problem of introducing noise. Algorithm-level methods may lead to overfitting of the model, and its classification effect is more dependent on the speci-fic dataset and classification task, which means they lack universality. In addition, how to deeply mine the differences in the distribution of data overlap areas, and how to effectively mine the differences between categories when the absolute number of minority samples is small, are also important chal-lenges in imbalanced classification. This paper proposes an imbalanced binary classification method using multi-label confidence comparisons based on contrastive learning. Different from the previous idea of directly learning its distribution characteristics from minority samples, combined with the idea of con-trastive learning, the classification task is redefined as the multi-label matching task by mining the deep features that can represent the commonality and difference between the neighboring samples. Multiple differentiated contrastive sample groups are obtained through random sampling in its neighbor sample pool for each sample. This sample is combined with its contrastive sample groups to form multiple sample-neighbor pairs as training samples in the multi-label matching task. The original dataset is mul-tiplied without introducing noise, laying a foundation for the effective mining of class differences when the absolute number of minority class samples is small. Based on the corresponding reconstruction error generated by Variational AutoEncoder (VAE), for sample-neighbor pairs, a multi-label matching loss between target samples and contrastive sample groups that integrates the idea of contrastive learning is designed. a robust classifier is obtained through simultaneous iterative learning of reconstruction error and multi-label matching loss, which can better mine the distribution differences of overlapping regions. In the testing phase, multiple different contrastive sample groups and the corresponding prediction results of the samples to be classified are obtained, which categories can be judged by integrating the pre-dictions of each group for reverse reasoning. Experimental results on 38 public datasets show that the method outperforms typical imbalanced classification methods in both F1-measure and G-mean.(c) 2022 Elsevier B.V. All rights reserved.
Keywords:
Imbalanced binary classification
Contrastive learning
Sample-neighbors pair
Variational auto-encoder
Multi-label confidence comparisons

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

B
beijing university of posts & telecommunications
Scholars:
1.4W
Papers: 1.2W
Citations: 9