arrow
返回

Large-scale incremental processing with MapReduce

delete2014-07-01
delete29
PRE
AI
D
Daewoo Lee *
J
Jin‐Soo Kim
S
Seungryoul Maeng
DOI:10.1016/j.future.2013.09.010delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
An important property of today's big data processing is that the same computation is often repeated on datasets evolving over time, such as web and social network data. While repeating full computation of the entire datasets is feasible with distributed computing frameworks such as Hadoop, it is obviously inefficient and wastes resources. In this paper, we present HadUP (Hadoop with Update Processing), a modified Hadoop architecture tailored to large-scale incremental processing with conventional MapReduce algorithms. Several approaches have been proposed to achieve a similar goal using task-level memoization. However, task-level memoization detects the change of datasets at a coarse-grained level, which often makes such approaches ineffective. Instead, HadUP detects and computes the change of datasets at a fine-grained level using a deduplication-based snapshot differential algorithm (D-SD) and update propagation. As a result, it provides high performance, especially in an environment where task-level memoization has no benefit. HadUP requires only a small amount of extra programming cost because it can reuse the code for the map and reduce functions of Hadoop. Therefore, the development of HadUP applications is quite easy. (C) 2013 Elsevier B.V. All rights reserved.
Keyword:
Big data processing
Incremental processing
MapReduce
Hadoop
Data deduplication
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
论文数:
6.8K
被引数:
2.3W

机构

S
sungkyunkwan university (skku)
学者数:
3.7W
论文数: 3.6W
被引数: 49