返回
Distributed multi-label feature selection using individual mutual information measures
DOI:10.1016/j.knosys.2019.105052.png)
摘要
En 中文
Multi-label learning generalizes traditional learning by allowing an instance to belong to multiple labels simultaneously. This causes multi-label data to be characterized by its large label space dimensionality and the dependencies among labels. These challenges have been addressed by feature selection techniques which improve the final model accuracy. However, the large number of features along with a large number of labels call for new approaches to manage data effectively and efficiently in distributed computing environments. This paper proposes a distributed model to compute a score that measures the quality of each feature with respect to multiple labels on Apache Spark. We propose two different approaches that study how to aggregate the mutual information of multiple labels: Euclidean Norm Maximization (ENM) and Geometric Mean Maximization (GMM). The former selects the features with the largest L-2-norm whereas the latter selects the features with the largest geometric mean. Experiments compare 9 distributed multi-label feature selection methods on 12 datasets and 12 metrics. Results validated through statistical analysis indicate that ENM is able to outperform the reference methods by maximizing the relevance while minimizing the redundancy of the selected features in constant selection time. (C) 2019 Elsevier B.V. All rights reserved.
Keyword:
Multi-label learning
Feature selection
Mutual information
Distributed computing
Apache spark
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
K
IF:
7.6
论文数:
1.3W
被引数:
4.5W
机构
引用论文
Joint multi-label classification and label correlations with missing labels and feature selection具有缺失标签和特征选择的联合多标签分类和标签相关性
Predicting protein structural classes for low-similarity sequences by evaluating different features通过评估不同特征预测低相似性序列的蛋白质结构课程
kNN-IS: An Iterative Spark-based design of the k-Nearest Neighbors classifier for big dataKnn-is: 基于Spark的大数据k近邻分类器迭代设计

