返回
Centralized vs. distributed feature selection methods based on data complexity measures
DOI:10.1016/j.knosys.2016.09.022.png)
摘要
En 中文
In the era of Big Data, many datasets have a common characteristic, the large number of features. As a result, selecting the relevant features and ignoring the irrelevant and redundant features has become indispensable. However, when dealing with large amounts of data, Most existing feature selection algorithms do not scale well, and their efficiency may significantly deteriorate to the point of becoming inapplicable. Moreover, data is often distributed in multiple locations, and it is not economic or legal to gather it in a single site. For these reasons, we propose a distributed approach for partitioned data using two techniques: horizontal (i.e. by samples) and vertical (i.e. by features). Unlike than existing procedures to combine the partial outputs obtained from each partition of data, we propose a merging process using the theoretical complexity of these feature subsets. The novel procedure tested in 11 datasets has proved to be useful, showing competitive results both in terms of runtime and classification accuracy. (C) 2016 Elsevier B.V. All rights reserved.
Keyword:
Distributed learning
Feature selection
Data complexity measures
Classification
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
K
IF:
7.6
论文数:
1.2W
被引数:
4.5W
机构
引用论文
Analysis of complexity indices for classification problems: Cancer gene expression data分类问题的复杂性指标分析: 癌症基因表达数据
NEUROCOMPUTING
IF6.5
Electrical Conductivity of Reproductive Tissue for Detection of Estrus in Dairy Cows用于检测奶牛发情的繁殖组织电导率
Physiological and behavioral responses to corticotropin-releasing factor administration: is CRF a mediator of anxiety or stress responses?促肾上腺皮质激素释放因子的生理和行为反应: CRF是焦虑或应激反应的中介物吗?
Bone marrow mesenchymal stem cells inhibit the response of naive and memory antigen-specific T cells to their cognate peptide
Blood
IF0
Learner excellence biased by data set selection: A case for data characterisation and artificial data sets
PATTERN RECOGNITION
IF7.6

