Return
Distributed Partial Label Learning for Missing Data Classification
DOI:10.3390/electronics14091770.png)
Abstract
En 中文
Distributed learning (DL), in which multiple nodes in an inner-connected network collaboratively induce a predictive model using their local data and some information communicated across neighboring nodes, has received significant research interest in recent years. Yet, it is challenging to achieve excellent performance in scenarios when training data instances have incomplete features and ambiguous labels. In such cases, it is essential to develop an efficient method to jointly perform the tasks of missing feature imputation and credible label recovery. Considering this, in this article, a distributed partial label missing data classification (dPMDC) algorithm is proposed. In the proposed algorithm, an integrated framework is formulated, which takes the ideas of both generative and discriminative learning into account. Firstly, by exploiting the weakly supervised information of ambiguous labels, a distributed probabilistic information-theoretic imputation method is designed to distributively fill in the missing features. Secondly, based on the imputed feature vectors, the classifier modeled by the random feature map of the chi(2) kernel function can be learned. Two iterative steps constitute the dPMDC algorithm, which can be used to handle dispersed, distributed data with partially missing features and ambiguous labels. Experiments on several datasets show the superiority of the suggested algorithm from many viewpoints.
Keywords:
distributed processing
partial label classification
missing data classification
random feature map of chi(2) kernel
Journal
IF:
2.6
Papers:
1.0W
Citations:
4.7W
Organization
No organization information available
Cited Papers
Effective Handling of Missing Values in Datasets for Classification Using Machine Learning Methods
Information
IF0

