Return
Self information-based feature selection for label distribution learning
DOI:10.1007/s13042-025-02771-1.png)
Abstract
En 中文
Label distribution learning (LDL) is an effective tool to process multi-label data where the label distribution is a probability distribution. Feature selection reduces data dimension, eliminates the impact of irrelevant features and enhances model performance. Rough set theory can be applied for feature selection in a LDL data. However, in most cases this theory only consider lower approximation when it is used to feature selection. In fact, uncertainty of information is related to both the upper and lower approximations. When measuring uncertainty, self information considers both the upper and lower approximations. This paper utilizes self information for feature selection in a LDL data. First of all, distance matrices in the feature space and the label space in a LDL data are constructed, respectively. Then, the upper and lower approximations in a LDL data are proposed. Subsequently, four types of self information (certain decision $$\alpha$$ -self information, possible decision $$\alpha$$ -self information, $$\alpha$$ -self information and relative $$\alpha$$ -self information) are defined to measure the uncertainty of a LDL data. Next, the best performance of self information: relative $$\alpha$$ -self information is selected by numerical analysis, a feature selection algorithm for a LDL data is designed using the selected self information. Finally, the designed algorithm is tested on 9 standard LDL datasets, and 6 indicators is used in experimental evaluation. The results demonstrate that the designed algorithm has the better performance of classification than 5 excellent feature selection algorithms.
Keywords:
Label distribution learning
Rough set theory
Self information
Feature selection
Uncertainty measurement
Neighborhood
Journal
IF:
2.7
Papers:
3.1K
Citations:
5.6K

