arrow
Return

Open world long-tailed data classification through active distribution optimization

delete2023-03-01
delete7
PRE
AI
M
Min Wang *
L
Lei Zhou
李谦 (Qian Li)
张安安 (Anan Zhang)
DOI:10.1016/j.eswa.2022.119054delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Real-world data exhibits a long-tailed label distribution, which leads to classification bias. Popular re-sampling or re-weighting methods usually require known category information. However, learning from long-tailed data with open categories is a challenging issue. In this paper, we propose an active distribution optimization algorithm (DALC) to handle the interesting issue. Through clustering, querying and classification iterations, we explore new categories and balance label distribution. For clustering, we present an exploration technique that adaptively obtains optimal data distribution with minimal total distance/cost. For each query, we design a critical instance selection strategy with the cluster information. For classification, we establish an ensemble model to continuously balance the label distribution. We conducted experiments on synthetic, benchmark and domain datasets. The results of the significance test verified the effectiveness of DALC and its superiority over state-of-the-art long-tailed data classification and open set classification algorithms.
Keywords:
Active learning
Cost-sensitive
Long-tailed distribution
Open set classification

Journal

Expert Systems with Applications cover
Expert Systems with Applications
IF:
7.5
Papers:
2.9W
Citations:
10.2W

Organization

S
Southwest Petroleum University
Scholars:
1.4W
Papers: 7.8K
Citations: 8.5K