返回
Machine Learning Feature Based Job Scheduling for Distributed Machine Learning Clusters
DOI:10.1109/TNET.2022.3190797.png)
摘要
En 中文
With the rapid proliferation of Machine Learning (ML) and Deep learning (DL) applications running on modern platforms, it is crucial to satisfy application performance requirements such as meeting deadline and ensuring accuracy. To this end, researchers have proposed several job schedulers for ML clusters. However, none of the previously proposed schedulers consider ML model parallelism, though it has been proposed as an approach to increase the efficiency of running large-scale ML and DL jobs. Thus, in this paper, we propose an ML job Feature based job Scheduling system (MLFS) for ML clusters running both data parallelism and model parallelism ML jobs. MLFS first uses a heuristic scheduling method that considers an ML job's spatial and temporal features to determine task priority for job queue ordering in order to improve job completion time (JCT) and accuracy performance. It uses the data from the heuristic scheduling method for training a deep reinforcement learning (RL) model. After the RL model is well trained, it then switches to the RL method to automatically make decisions on job scheduling. In addition, MLFS has a system load control method that selects tasks from overloaded servers to move to underloaded servers based on task priority, and also intelligently removes the tasks that generate little or no improvement on the desired accuracy performance when the system is overloaded to improve JCT and accuracy by job deadline. Furthermore, we propose Optimal ML iteration stopping method that determines the proper time to stop training ML model when this model reaches the minimum loss value. Our real experiments and large-scale simulation based on real trace show that MLFS reduces JCT by up to 53% and makespan by up to 52%, and improves accuracy by up to 64% when compared with existing ML job schedulers. We also open sourced our code.
Keyword:
Task analysis
Parallel processing
Servers
Scheduling
Data models
Training
Graphics processing units
Machine learning
resource management
job scheduling
期刊
I
IF:
3.6
论文数:
4.4K
被引数:
9.5K
机构
引用论文
Succinylated copper, zinc superoxide dismutase. A novel approach to the problem of active subunits琥珀酰化铜,锌超氧化物歧化酶。一种解决活性亚基问题的新方法
Biochemistry
IF0
Adherence of sickle erythrocytes to vascular endothelial cells: requirement for both cell membrane changes and plasma factors
Blood
IF0

