arrow
返回

Big data classification using heterogeneous ensemble classifiers in Apache Spark based on MapReduce paradigm

delete2021-11-01
delete14
PRE
AI
H
Hamidreza Kadkhodaei
M
Mehdi Dehghan
DOI:10.1016/j.eswa.2021.115369delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In this era of big data, processing large scale data efficiently and accurately has become a challenging problem. Ensemble classification is a type of supervised learning that uses multiple experts to generate the final output. It provides a way to classify data more accurately. As a result of using multiple classifiers, they are often more complicated than single classifiers, especially for big data problems. Apache Spark is a unified analytics engine for big data processing which provides a scalable framework to analyze the data. In this paper, we first extend our previous work and design a distributed heterogeneous ensemble classifier inspired by the boosting approach, which is capable of dealing with big datasets. Using heterogeneous classifiers makes it possible to have more diverse classifiers, and consequently, a more accurate classifier is obtained. Then, we present the Spark version of the proposed approach to speed up our heterogeneous ensemble classifier using the MapReduce paradigm. In order to evaluate our approach, we have applied it to seven big datasets. Extensive experimental results indicate the superiority of the proposed method over the existing ensemble algorithms implemented by Spark MLlib in terms of the classification accuracy, performance, and scalability.
Keyword:
Ensemble classifier
Boosting
MapReduce
Big data
Apache Spark
Apache Hadoop
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Expert Systems with Applications 封面图
Expert Systems with Applications
IF:
7.5
论文数:
3.0W
被引数:
10.2W

机构

I
Islamic Azad University
学者数:
4.0W
论文数: 3.3W
被引数: 9.8K
A
Amirkabir University of Technology
学者数:
1.1W
论文数: 1.1W
被引数: 1.0W
引用论文

引用论文

err分享
err收藏
Classifier Subset Selection to construct multi-classifiers by means of estimation of distribution algorithms
err2015-06-01
err31
errOAAI
errMendialdua, Inigo; Arruti, Andoni; Jauregi, Ekaitz; Lazkano, Elena; Sierra, Basilio
err分享
err收藏
err分享
err收藏
The strength of weak learnability
err1990-06-01
err0
errOAAI
errRobert E. Schapire
err分享
err收藏
学者 查看更多内容