arrow
返回

Machine Learning Feature Based Job Scheduling for Distributed Machine Learning Clusters

delete2023-02-01
delete3
PRE
AI
H
Haoyu Wang
Z
Zetian Liu
H
Haiying Shen *
DOI:10.1109/TNET.2022.3190797delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
With the rapid proliferation of Machine Learning (ML) and Deep learning (DL) applications running on modern platforms, it is crucial to satisfy application performance requirements such as meeting deadline and ensuring accuracy. To this end, researchers have proposed several job schedulers for ML clusters. However, none of the previously proposed schedulers consider ML model parallelism, though it has been proposed as an approach to increase the efficiency of running large-scale ML and DL jobs. Thus, in this paper, we propose an ML job Feature based job Scheduling system (MLFS) for ML clusters running both data parallelism and model parallelism ML jobs. MLFS first uses a heuristic scheduling method that considers an ML job's spatial and temporal features to determine task priority for job queue ordering in order to improve job completion time (JCT) and accuracy performance. It uses the data from the heuristic scheduling method for training a deep reinforcement learning (RL) model. After the RL model is well trained, it then switches to the RL method to automatically make decisions on job scheduling. In addition, MLFS has a system load control method that selects tasks from overloaded servers to move to underloaded servers based on task priority, and also intelligently removes the tasks that generate little or no improvement on the desired accuracy performance when the system is overloaded to improve JCT and accuracy by job deadline. Furthermore, we propose Optimal ML iteration stopping method that determines the proper time to stop training ML model when this model reaches the minimum loss value. Our real experiments and large-scale simulation based on real trace show that MLFS reduces JCT by up to 53% and makespan by up to 52%, and improves accuracy by up to 64% when compared with existing ML job schedulers. We also open sourced our code.
Keyword:
Task analysis
Parallel processing
Servers
Scheduling
Data models
Training
Graphics processing units
Machine learning
resource management
job scheduling

期刊

I
IEEE-ACM Transactions on Networking
IF:
3.6
论文数:
4.4K
被引数:
9.5K

机构

U
University of Virginia
学者数:
3.0W
论文数: 2.7W
被引数: 4.1W
引用论文

引用论文

A Practical Guide to Information Analysis of Spike Trains
err2003-01-01
err0
PREAI
errGianni Pola; Simon R. Schultz; Rasmus S. Petersen; Stefano Panzeri
err分享
err收藏
An Exploratory Survey ofDeqiSensation from the Views and Experiences of Chinese Patients and Acupuncturists
err2013-01-01
err0
errOAAI
errHong-Wen Yuan; Liang-Xiao Ma; Peng Zhang; Chi Lin; Dan-Dan Qi; Jing Li; Si-Yuan Xin; Ni-Juan Hu; Chun-Hua Li; Yu-Qi Liu; Jie Hao; Jie-Ping Xie; Hai Cui; Jiang Zhu
err分享
err收藏
Recombinant humanized anti-PD-1 monoclonal antibody toripalimab in patients with metastatic urothelial carcinoma: Results of an open-label phase II clinical study Polaris-03.
err2020-05-20
err0
PREAI
errXinan Sheng; Haige Chen; Bin Hu; Xudong Yao; Ziling Liu; Xin Yao; Hongqian Guo; Yi Hu; Zhigang Ji; Hong Luo; Benkang Shi; Jiyan Liu; Jin WU; Fangjian Zhou; Zhisong He; Jinhai Fan; Yiran Huang; Jun Guo
err分享
err收藏
Assessing the generalizability of eye dominance across binocular rivalry, onset rivalry, and continuous flash suppression
err2018-06-21
err0
errOAAI
errYun Ding; Marnix Naber; Surya Gayet; Stefan Van der Stigchel; Chris L. E. Paffen
err分享
err收藏
学者 查看更多内容