Return
Self-Supervised Learning for Labeling Unlabeled Imbalanced Data using Generalized Label Distribution
DOI:10.1016/j.jfranklin.2025.108230.png)
Abstract
En 中文
Motivated to label unlabeled data while using minimum prior information, a new self-supervised learning approach is proposed in this study where the specific initial data labels are not compulsorily required yet the genuine labels can be generated. The new self-supervised learning approach is named CCCV as it is with four steps: clustering (C) for generating initial labels using the generalized label distribution (GLD), the calculation (C) of the data-label degree, correction (C) if the labels are misassigned according to the data-label degree, and validation (V). Compared with supervised learning, CCCV requires only unlabeled data. Compared with unsupervised learning, CCCV can identify genuine labels. Compared with semi- or weak-supervised learning approaches, CCCV has a much-relaxed requirement on the accuracy of initial data labels owing to GLD. The following conclusions are drawn from the case study results on 11 UCI benchmarks when their labels are not used in clustering, calculation, and correction (only used for validation as the gold standard). (1) The accuracy of 11 benchmarks is significantly improved compared with the initial clustering approaches when data labels are not used; the accuracy is even competitive compared with that of classification approaches when data labels are used. (2) The data-label degree can effectively help in correcting the mis-assigned labels in all cases. (3) The generalization ability of CCCV over the initial step of clustering is validated as six clustering approaches are tested and consistent results are produced.
Journal
J
IF:
4.2
Papers:
822
Citations:
0

