返回
Toward Efficient Online Scheduling for Distributed Machine Learning Systems
DOI:10.1109/TNSE.2021.3104513.png)
摘要
En 中文
Recent years have witnessed a rapid growth of distributed machine learning (ML) frameworks, which exploit the massive parallelism of computing clusters to expedite ML training. However, the proliferation of distributed ML frameworks also introduces many unique technical challenges in computing system design and optimization. In a networked computing cluster that supports a large number of training jobs, a key question is how to design efficient scheduling algorithms to allocate workers and parameter servers across different machines to minimize the overall training time. Toward this end, in this paper, we develop an online scheduling algorithm that jointly optimizes resource allocation and locality decisions. Our main contributions are three-fold: i) We develop a new analytical model that considers both resource allocation and locality; ii) Based on an equivalent reformulation and observations on the worker-parameter server locality configurations, we transform the problem into a mixed packing and covering integer program, which enables approximation algorithm design; iii) We propose a meticulously designed approximation algorithm based on randomized rounding and rigorously analyze its performance. Collectively, our results contribute to the state of the art of distributed ML system optimization and algorithm design.
Keyword:
Servers
Training
Optimization
Scheduling algorithms
Resource management
Approximation algorithms
Heuristic algorithms
Online resource scheduling
distributed machine learning
approximation algorithm
期刊
I
IF:
7.9
论文数:
2.6K
被引数:
10.0K
机构
引用论文
Succinylated copper, zinc superoxide dismutase. A novel approach to the problem of active subunits琥珀酰化铜,锌超氧化物歧化酶。一种解决活性亚基问题的新方法
Biochemistry
IF0
Adherence of sickle erythrocytes to vascular endothelial cells: requirement for both cell membrane changes and plasma factors
Blood
IF0

