arrow
返回

Achieving Load Balance for Parallel Data Access on Distributed File Systems

delete2018-03-01
delete26
delete
OA
AI
D
Dan Huang
D
Dezhi Han *
J
Jun Wang *
X
Xunchao Chen
X
Xuhong Zhang
J
Jian Zhou
M
Mao Ye
DOI:10.1109/TC.2017.2749229delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
The distributed file system, HDFS, is widely deployed as the bedrock for many parallel big data analysis. However, when running multiple parallel applications over the shared file system, the data requests from different processes/executors will unfortunately be served in a surprisingly imbalanced fashion on the distributed storage servers. These imbalanced access patterns among storage nodes are caused because a). unlike conventional parallel file system using striping policies to evenly distribute data among storage nodes, data-intensive file system such as HDFS store each data unit, referred to as chunk file, with several copies based on a relative random policy, which can result in an uneven data distribution among storage nodes; b). based on the data retrieval policy in HDFS, the more data a storage node contains, the higher probability the storage node could be selected to serve the data. Therefore, on the nodes serving multiple chunk files, the data requests from different processes/executors will compete for shared resources such as hard disk head and network bandwidth, resulting in a degraded I/O performance. In this paper, we first conduct a complete analysis on how remote and imbalanced read/write patterns occur and how they are affected by the size of the cluster. We then propose novel methods, referred to as Opass, to optimize parallel data reads, as well as to reduce the imbalance of parallel writes on distributed file systems. Our proposed methods can benefit parallel data-intensive analysis with various parallel data access strategies. Opass adopts new matching-based algorithms to match processes to data so as to compute the maximum degree of data locality and balanced data access. Furthermore, to reduce the imbalance of parallel writes, Opass employs a heatmap for monitoring the I/O statuses of storage nodes and performs HM-LRU policy to select a local optimal storage node for serving write requests. Experiments are conducted on PRObE's Marmot 128-node cluster testbed and the results from both benchmark and well-known parallel applications show the performance benefits and scalability of Opass.
Keyword:
Parallel data access
distributed file systems
HDFS
bipartite matching
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Computers 封面图
IEEE Transactions on Computers
IF:
3.8
论文数:
5.4K
被引数:
9.8K

机构

State University System of Florida 封面图
State University System of Florida
学者数:
12.8W
论文数: 10.9W
被引数: 130
U
University of Central Florida
学者数:
8.7K
论文数: 6.8K
被引数: 1.4W
引用论文

引用论文

DNA Binding to the Silica Surface
err2015-06-02
err0
PREAI
errBobo Shi; Yun Kyung Shin; Ali A. Hassanali; Sherwin J. Singer
err分享
err收藏
Kinetic resolution strategies using non-enzymatic catalysts
err2003-06-01
err0
PREAI
errDiane E.J.E. Robinson; Steven D. Bull
err分享
err收藏
err分享
err收藏
err分享
err收藏
Regulation Inside Government: Where New Public Management Meets the Audit Explosion
err1998-04-01
err0
PREAI
errChristopher Hood; Oliver James; George Jones; Colin Scott; Tony Travers
err分享
err收藏
学者 查看更多内容