Return
Multi-label feature selection based on positive sample information weighting
DOI:10.1016/j.knosys.2025.114373.png)
Abstract
En 中文
Multi-label feature selection has garnered significant attention due to its capacity to reduce data dimensionality while effectively eliminating noise and irrelevant features. Information-theoretic methods are predominant in this domain, as they aim to quantify the relevance and redundancy among features, as well as between features and labels. However, existing information-theoretic methods frequently disregard the inherent distributional differences between positive and negative samples when assessing relevance and redundancy among variables. In multi-label data sets, positive samples associated with the same label tend to be more concentrated in the feature space, whereas negative samples exhibit a more dispersed distribution. This disparity is particularly evident in datasets characterized by label sparsity. The dispersed distribution of feature values for negative samples often results in reduced discriminative power and increased susceptibility to noise. To mitigate this limitation, we propose a novel feature selection method termed Multi-label Feature Selection based on Positive Sample Information Weighting (PSIWFS). PSIWFS assigns higher weights to features that demonstrate strong correlations with positive samples during the feature selection process, thereby enhancing the identification and prioritization of the less frequent positive samples. Experimental evaluations conducted on 14 datasets with 7 comparison methods underscore the superior classification performance of the proposed method.

