返回
MapReduce based improved quick reduct algorithm with granular refinement using vertical partitioning scheme
DOI:10.1016/j.knosys.2019.105104.png)
摘要
En 中文
In the last few decades, rough sets have evolved to become an essential technology for feature subset selection by way of reduct computation in categorical decision systems. In recent years with the proliferation of MapReduce for distributed/parallel algorithms, several scalable reduct computation algorithms have been developed in this field for large-scale decision systems using MapReduce. The existing MapReduce based reduct computation approaches use horizontal partitioning (division in object space) of the dataset into the nodes of the cluster, requiring a complicated shuffle and sort phase. In this work, we propose an algorithm MR_IQRA_VP which is designed using vertical partitioning (division in attribute space) of the dataset with a simplified shuffle and sort phase of the MapReduce framework. MR_IQRA_VP is a distributed/parallel implementation of the Improved Quick Reduct Algorithm (IQRA_IG) and is implemented using iterative MapReduce framework of Apache Spark. We have done an extensive comparative study through experimentation on benchmark decision systems using existing horizontal partitioning based reduct computation algorithms. Through experimental analysis, along with theoretical validation, we have established that MR_IQRA_VP is suitable and scalable to datasets of larger size attribute space and moderate object space prevalent in the areas of Bioinformatics and Web mining. (C) 2019 Elsevier B.V. All rights reserved.
Keyword:
Rough sets
MapReduce
Apache spark
Reduct
Horizontal partitioning
Vertical partitioning
Feature subset selection
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
K
IF:
7.6
论文数:
1.3W
被引数:
4.5W
机构
引用论文
Late gadolinium enhancement on cardiac magnetic resonance combined with 123I- metaiodobenzylguanidine scintigraphy strongly predicts long-term clinical outcome in patients with dilated cardiomyopathy
PLOS ONE
IF0
An Information Theory-Based Feature Selection Framework for Big Data Under Apache SparkApache Spark下基于信息论的大数据特征选择框架
An efficient accelerator for attribute reduction from incomplete data in rough set framework
PATTERN RECOGNITION
IF7.6

