arrow
Return

An Information Theory-Based Feature Selection Framework for Big Data Under Apache Spark

delete2018-09-01
delete49
PRE
AI
S
Sergio Ramírez‐Gallego *
V
Verónica Bolón‐Canedo
J
José M. Benítez
A
Amparo Alonso‐Betanzos
F
Francisco Herrera
DOI:10.1109/TSMC.2017.2670926delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
With the advent of extremely high dimensional datasets, dimensionality reduction techniques are becoming mandatory. Of the many techniques available, feature selection (FS) is of growing interest for its ability to identify both relevant features and frequently repeated instances in huge datasets. We aim to demonstrate that standard FS methods can be parallelized in big data platforms like Apache Spark so as to boost both performance and accuracy. We propose a distributed implementation of a generic FS framework that includes a broad group of well-known information theory-based methods. Experimental results for a broad set of real-world datasets show that our distributed framework is capable of rapidly dealing with ultrahigh-dimensional datasets as well as those with a huge number of samples, outperforming the sequential version in all the cases studied.
Keywords:
Apache spark
big data
feature selection (FS)
filtering methods
high-dimensional
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Cybernetics cover
IEEE Transactions on Cybernetics
IF:
10.5
Papers:
1.1W
Citations:
5.0W

Organization

U
Universidade da Coruna
Scholars:
6.6K
Papers: 5.7K
Citations: 11
U
University of Granada
Scholars:
2.3W
Papers: 1.9W
Citations: 24