arrow
Return

Stargazer: Toward efficient data analytics scheduling via task completion time inference

delete2021-06-01
delete3
PRE
AI
H
Haizhou Du *
K
Keke Zhang
Q
Qiao Xiang
DOI:10.1016/j.compeleceng.2021.107092delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The fundamental challenge of data analytics scheduling is the heterogeneity of both data analytics jobs and resources. Although many scheduling solutions have been developed to improve the efficiency of data analytics frameworks (e.g., Spark), they either (1) focus on the scheduling of a single type of resource, without considering the coordination between different resources; or (2) schedule multiple resources by factoring in limited information about analytics jobs without considering the heterogeneity of resources. This paper presents Stargazer, a novel, efficient system that tackles diversity data analytics jobs on heterogeneous cluster by inferring the completion times of their decomposed tasks. Specifically, Stargazer adopts a deep learning model, which takes into considerations multiple key factors of diversity data analytics jobs and heterogeneous resources, to accurately infer the completion time of different tasks. A prototype of Stargazer is fully implemented in the Spark framework. Extensive experiments show that Stargazer can reduce the average job completion time by 21% and improve average performance by 20%, while incurring little overhead.
Keywords:
Spark scheduling optimization
Delay scheduling
Computation complexity
Deep learning
Data locality

Journal

C
Computers and Electrical Engineering
IF:
4.9
Papers:
6.7K
Citations:
1.3W

Organization

T
tongji university
Scholars:
7.7W
Papers: 5.9W
Citations: 98
S
Shanghai University of Electric Power
Scholars:
5.2K
Papers: 3.4K
Citations: 4.9K