arrow
返回

Centralized vs. distributed feature selection methods based on data complexity measures

delete2017-02-01
delete46
PRE
AI
B
Bolon-Canedo, V.
A
Amparo Alonso‐Betanzos
DOI:10.1016/j.knosys.2016.09.022delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In the era of Big Data, many datasets have a common characteristic, the large number of features. As a result, selecting the relevant features and ignoring the irrelevant and redundant features has become indispensable. However, when dealing with large amounts of data, Most existing feature selection algorithms do not scale well, and their efficiency may significantly deteriorate to the point of becoming inapplicable. Moreover, data is often distributed in multiple locations, and it is not economic or legal to gather it in a single site. For these reasons, we propose a distributed approach for partitioned data using two techniques: horizontal (i.e. by samples) and vertical (i.e. by features). Unlike than existing procedures to combine the partial outputs obtained from each partition of data, we propose a merging process using the theoretical complexity of these feature subsets. The novel procedure tested in 11 datasets has proved to be useful, showing competitive results both in terms of runtime and classification accuracy. (C) 2016 Elsevier B.V. All rights reserved.
Keyword:
Distributed learning
Feature selection
Data complexity measures
Classification
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

K
Knowledge-Based Systems
IF:
7.6
论文数:
1.2W
被引数:
4.5W

机构

U
Universidade da Coruna
学者数:
6.6K
论文数: 5.7K
被引数: 11
引用论文

引用论文

The electroneutrality approximation in electrochemistry电化学中的电中性近似
err2011-02-22
err0
PREAI
errEdmund J. F. Dickinson; Juan G. Limon-Petersen; Richard G. Compton
err分享
err收藏
err分享
err收藏
err分享
err收藏
Bone marrow mesenchymal stem cells inhibit the response of naive and memory antigen-specific T cells to their cognate peptide
err2003-05-01
err0
errOAAI
errMauro Krampera; Sarah Glennie; Julian Dyson; Diane Scott; Ruthline Laylor; Elizabeth Simpson; Francesco Dazzi
err分享
err收藏
Learner excellence biased by data set selection: A case for data characterisation and artificial data sets
err2013-03-01
err37
PREAI
errMacia, Nuria; Bernado-Mansilla, Ester; Orriols-Puig, Albert; Ho, Tin Kam
err分享
err收藏
Recent advances and emerging challenges of feature selection in the context of big data
err2015-09-01
err199
PREAI
errBolon-Canedo, V.; Sanchez-Marono, N.; Alonso-Betanzos, A.
err分享
err收藏
学者 查看更多内容