arrow
Return

Machine Learning Feature Based Job Scheduling for Distributed Machine Learning Clusters

delete2023-02-01
delete3
PRE
AI
H
Haoyu Wang
Z
Zetian Liu
H
Haiying Shen *
DOI:10.1109/TNET.2022.3190797delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
With the rapid proliferation of Machine Learning (ML) and Deep learning (DL) applications running on modern platforms, it is crucial to satisfy application performance requirements such as meeting deadline and ensuring accuracy. To this end, researchers have proposed several job schedulers for ML clusters. However, none of the previously proposed schedulers consider ML model parallelism, though it has been proposed as an approach to increase the efficiency of running large-scale ML and DL jobs. Thus, in this paper, we propose an ML job Feature based job Scheduling system (MLFS) for ML clusters running both data parallelism and model parallelism ML jobs. MLFS first uses a heuristic scheduling method that considers an ML job's spatial and temporal features to determine task priority for job queue ordering in order to improve job completion time (JCT) and accuracy performance. It uses the data from the heuristic scheduling method for training a deep reinforcement learning (RL) model. After the RL model is well trained, it then switches to the RL method to automatically make decisions on job scheduling. In addition, MLFS has a system load control method that selects tasks from overloaded servers to move to underloaded servers based on task priority, and also intelligently removes the tasks that generate little or no improvement on the desired accuracy performance when the system is overloaded to improve JCT and accuracy by job deadline. Furthermore, we propose Optimal ML iteration stopping method that determines the proper time to stop training ML model when this model reaches the minimum loss value. Our real experiments and large-scale simulation based on real trace show that MLFS reduces JCT by up to 53% and makespan by up to 52%, and improves accuracy by up to 64% when compared with existing ML job schedulers. We also open sourced our code.
Keywords:
Task analysis
Parallel processing
Servers
Scheduling
Data models
Training
Graphics processing units
Machine learning
resource management
job scheduling

Journal

I
IEEE-ACM Transactions on Networking
IF:
3.6
Papers:
4.4K
Citations:
9.5K

Organization

U
University of Virginia
Scholars:
3.0W
Papers: 2.7W
Citations: 4.1W
Cited Papers

Cited Papers

Succinylated copper, zinc superoxide dismutase. A novel approach to the problem of active subunits
err2002-05-01
err0
PREAI
errFranco Marmocchi; Irene Mavelli; Adelio Rigo; Roberto Stevanato; Francesco Bossa; Giuseppe Rotilio
errShare
errSave
A Practical Guide to Information Analysis of Spike Trains
err2003-01-01
err0
PREAI
errGianni Pola; Simon R. Schultz; Rasmus S. Petersen; Stefano Panzeri
errShare
errSave
An Exploratory Survey ofDeqiSensation from the Views and Experiences of Chinese Patients and Acupuncturists
err2013-01-01
err0
errOAAI
errHong-Wen Yuan; Liang-Xiao Ma; Peng Zhang; Chi Lin; Dan-Dan Qi; Jing Li; Si-Yuan Xin; Ni-Juan Hu; Chun-Hua Li; Yu-Qi Liu; Jie Hao; Jie-Ping Xie; Hai Cui; Jiang Zhu
errShare
errSave
Recombinant humanized anti-PD-1 monoclonal antibody toripalimab in patients with metastatic urothelial carcinoma: Results of an open-label phase II clinical study Polaris-03.
err2020-05-20
err0
PREAI
errXinan Sheng; Haige Chen; Bin Hu; Xudong Yao; Ziling Liu; Xin Yao; Hongqian Guo; Yi Hu; Zhigang Ji; Hong Luo; Benkang Shi; Jiyan Liu; Jin WU; Fangjian Zhou; Zhisong He; Jinhai Fan; Yiran Huang; Jun Guo
errShare
errSave
Assessing the generalizability of eye dominance across binocular rivalry, onset rivalry, and continuous flash suppression
err2018-06-21
err0
errOAAI
errYun Ding; Marnix Naber; Surya Gayet; Stefan Van der Stigchel; Chris L. E. Paffen
errShare
errSave
researcher View more