arrow
返回

CloudFlow: A data-aware programming model for cloud workflow applications on modern HPC systems

delete2015-10-01
delete9
PRE
AI
F
Fan Zhang *
Q
Qutaibah Malluhi
T
Tamer Elsayed
S
Samee U. Khan
李克勤 封面图
李克勤 (Keqin Li)
A
Albert Y. Zomaya
DOI:10.1016/j.future.2014.10.028delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Traditional High-Performance Computing (HPC) based big-data applications are usually constrained by having to move large amount of data to compute facilities for real-time processing purpose. Modern HPC systems, represented by High-Throughput Computing (HTC) and Many-Task Computing (MTC) platforms, on the other hand, intend to achieve the long-held dream of moving compute to data instead. This kind of data-aware scheduling, typically represented by Hadoop MapReduce, has been successfully implemented in its Map Phase, whereby each Map Task is sent out to the compute node where the corresponding input data chunk is located. However, Hadoop MapReduce limits itself to a one-map-to-one-reduce framework, leading to difficulties for handling complex logics, such as pipelines or workflows. Meanwhile, it lacks built-in support and optimization when the input datasets are shared among multiple applications and/or jobs. The performance can be improved significantly when the knowledge of the shared and frequently accessed data is taken into scheduling decisions. To enhance the capability of managing workflow in modern HPC system, this paper presents CloudFlow, a Hadoop MapReduce based programming model for cloud workflow applications. CloudFlow is built on top of MapReduce, which is proposed not only being data aware, but also shared-data aware. It identifies the most frequently shared data, from both task-level and job-level, replicates them to each compute node for data locality purposes. It also supports user-defined multiple Map- and Reduce functions, allowing users to orchestrate the required data-flow logic. Mathematically, we prove the correctness of the whole scheduling framework by performing theoretical analysis. Further more, experimental evaluation also shows that the execution runtime speedup exceeds 4X compared to traditional MapReduce implementation with a manageable time overhead. (C) 2014 Elsevier B.V. All rights reserved.
Keyword:
Concurrency
Data aware
MapReduce
HPC
Programming model
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
论文数:
6.8K
被引数:
2.3W

机构

N
north dakota state university fargo
学者数:
5.5K
论文数: 4.9K
被引数: 6
S
state university of new york (suny) system
学者数:
6.5W
论文数: 5.8W
被引数: 65
SUNY New Paltz 封面图
SUNY New Paltz
学者数:
169
论文数: 199
被引数: 324
Q
Qatar University
学者数:
8.9K
论文数: 9.0K
被引数: 16
学者 查看更多机构
引用论文

引用论文

Mechanism of subwavelength imaging with bilayered magnetic metamaterials: Theory and experiment
err2007-04-03
err0
errOAAI
errO. Sydoruk; M. Shamonin; A. Radkovskaya; O. Zhuromskyy; E. Shamonina; R. Trautner; C. J. Stevens; G. Faulkner; D. J. Edwards; L. Solymar
err分享
err收藏
err分享
err收藏
Sentence processing is uniquely human
err2003-07-01
err0
PREAI
errKuniyoshi L. Sakai; Fumitaka Homae; Ryuichiro Hashimoto
err分享
err收藏
BIOMASS DYNAMICS IN AMAZONIAN FOREST FRAGMENTS
err2004-08-01
err0
PREAI
errHenrique E. M. Nascimento; William F. Laurance
err分享
err收藏
学者 查看更多内容