arrow
Return

A Multidimensional Communication Scheduling Method for Hybrid Parallel DNN Training

delete2024-08-01
delete0
PRE
AI
S
Shengwei Li
K
Kai Lü
Z
Zhiquan Lai
W
Weijie Liu
K
Keshi Ge
D
Dongsheng Li *
DOI:10.1109/TPDS.2024.3406420delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The transformer-based deep neural network (DNN) models have shown considerable success across diverse tasks, prompting widespread adoption of distributed training methods such as data parallelism and pipeline parallelism. With the increasing parameter number, hybrid parallel training becomes imperative to scale training. The primary bottleneck in scaling remains the communication overhead. The communication scheduling technique, emphasizing the overlap of communication with computation, has demonstrated its benefits in scaling. However, most existing works focus on data parallelism, overlooking the nuances of hybrid parallel training. In this paper, we propose TriRace, an efficient communication scheduling framework for accelerating communications in hybrid parallel training of asynchronous pipeline parallelism and data parallelism. To achieve effective computation-communication overlap, TriRace introduces 3D communication scheduling, which adeptly leverages data dependencies between communication and computations, efficiently scheduling AllReduce communication, sparse communication, and peer-to-peer communication in hybrid parallel training. To avoid possible communication contentions, TriRace also incorporates a topology-aware runtime which optimizes the execution of communication operations by considering ongoing communication operations and real-time network status. We have implemented a prototype of TriRace based on PyTorch and Pipedream-2BW, and conducted comprehensive evaluations with three representative baselines. Experimental results show that TriRace achieves up to 1.07-1.45x speedup compared to the state-of-the-art pipeline parallelism training baseline Pipedream-2BW, and 1.24-1.81x speedup compared to the Megatron.
Keywords:
Training
Computational modeling
Processor scheduling
Pipelines
Data models
Transformers
Pipeline processing
Distributed training
deep learning
hybrid parallelism
communication scheduling

Journal

IEEE Transactions on Parallel and Distributed Systems cover
IEEE Transactions on Parallel and Distributed Systems
IF:
6
Papers:
5.2K
Citations:
1.1W

Organization

N
national university of defense technology - china
Scholars:
1.8W
Papers: 1.4W
Citations: 9