arrow
Return

A fast parallel attribute reduction algorithm using Apache Spark

delete2021-01-01
delete15
PRE
AI
L
Linzi Yin
Z
Zhaohui Jiang
X
Xuemei Xu
DOI:10.1016/j.knosys.2020.106582delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Effective and fast attribute reduction algorithm on high-dimensional dataset is one of the most important issues of big data, and several parallel attribute reduction algorithms were implemented by using MapReduce. However, MapReduce is not suitable for iterative computing, which causes low calculation efficiency in many cases. In this paper, we proposed a novel parallel attribute reduction algorithm by considering the new generation distributed computing framework Apache Spark. First, the core attribute decision strategy is proposed to replace the traditional attribute significance calculation, and the number of iterations is reduced from vertical bar C vertical bar vertical bar R vertical bar-vertical bar R vertical bar(2)/2+vertical bar R vertical bar/2 to vertical bar C vertical bar (vertical bar C vertical bar represents the number of condition attributes and vertical bar R vertical bar represents the number of attributes in the reduct result). Furthermore, for high-dimensional datasets, we designed a batch processing strategy to reduce the number of iterations exponentially. Second, the proposed algorithm was speeded up with three techniques, including: (1) the network data transmission is minimized based on the localized operation; (2) a single cache iteration method is suggested to reduce disk I/O cost; (3) some calculations are skipped by an interruption strategy. In the experimental analysis, we succeeded with various types of real big datasets and random datasets in a real distributed computing environment and compared with the classic MapReduce-based parallel attribute reduction algorithm PAAR_PR in various aspects. Experimental conclusions proved that the computing efficiency of our algorithm has been improved by more than 98% compared to the classic parallel attribute reduction algorithm PAAR_PR. (C) 2020 Elsevier B.V. All rights reserved.
Keywords:
Rough sets
Big data
Parallel algorithm
Attribute reduction
Apache Spark
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

C
Central South University
Scholars:
10.0W
Papers: 7.2W
Citations: 10.9W