返回
An Enhanced Data-Locality-Aware Task Scheduling Algorithm for Hadoop Applications
DOI:10.1109/JSYST.2017.2764481.png)
摘要
En 中文
In general, Hadoop improves the task scheduling performance by determining data locality based on the location in which the input splits and MapTask are executed. However, if an input split consists of multiple data blocks that are distributed and stored in different nodes, this data location method fails to cope with the degradation in processing performance due to the increased frequency of data block copying. We propose a task scheduling algorithm that solves this issue by defining a method to classify data locality taking into account the location of all data blocks that comprise an input split, categorizing tasks based on the defined method, and sequentially assigning tasks according to a given priority. This study measures the performance of the proposed algorithm through a comparison of the total processing time, MapTask performance time, and data block copying frequency between the proposed algorithm and Hadoop's default task scheduling algorithm. The test results show that the proposed algorithm improved the total processing time by up to 25% and the data block copying frequency by up to 28%, when compared to the default algorithm.
Keyword:
Data locality
Hadoop distributed file system (HDFS)
MapReduce
task scheduling
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
I
IF:
2.4
论文数:
4.5K
被引数:
387
机构
引用论文
MEDULLARY SPONGE KIDNEY DISEASE AND RECURRENT NEPHROLITHIASIS: IS DISTAL RENAL TUBULAR ACIDOSIS THE LAST PIECE OF THE PUZZLE?髓质海绵肾疾病与复发性肾结石:远端肾小管酸中毒是最后一块拼图吗?
Effect of stepwise protonation of an N-containing ligand on the formation of metal–organic salts and coordination complexes in the solid state
CrystEngComm
IF0
DOUBLE TROUBLE: A RARE CASE OF DUAL-POSITIVE ANTI-GBM AND ANCA VASCULITIS双重麻烦:一例罕见的GBM和ANCA双重阳性血管炎病例

