arrow
返回

Accelerating Distributed DNN Training via Transport Layer Scheduling

delete2023-05-01
delete1
PRE
AI
Q
Qingyang Duan
C
Chao Peng
Z
Zeqin Wang
徐
徐跃东 (Yuedong Xu) *
J
Jun Wu
J
John C. S. Lui
DOI:10.1109/TPDS.2023.3250462delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Communication scheduling is crucial to accelerate the training of large deep learning models, in which the transmission order of layer-wise deep neural network (DNN) tensors is determined for a better computation-communication overlap. Prior approaches adopt user-level tensor partitioning to enhance the priority scheduling with finer granularity. However, a startup time slot inserted before every tensor partition will neutralize this scheduling gain. Tuning hyper-parameters for tensor partitioning is difficult, especially when the network bandwidth is shared or time-varying in multi-tenant clusters. In this article, we propose Mercury, a simple transport layer scheduler that moves the priority scheduling to the transport layer at the packet granularity. The packets with the highest priority in the Mercury buffer will be transmitted first. Mercury achieves the near-optimal overlapping between communication and computation. It also leverages the immediate aggregation at the transport layer to enable the full overlapping of gradient push and pull. We implement Mercury in MXNet and conduct comprehensive experiments on five popular DNN models in various environments. Mercury can well adapt to dynamic communication and computation resources. Experiments show that Mercury accelerates the training by up to 130% compared to the classical PS architecture, and 104% compared to state-of-the-art tensor partitioning methods.
Keyword:
Tensors
Training
Processor scheduling
Computer architecture
Computational modeling
Bandwidth
Servers
Computation-communication overlap
distri- buted machine learning
parameter server
transport layer scheduling

期刊

IEEE Transactions on Parallel and Distributed Systems 封面图
IEEE Transactions on Parallel and Distributed Systems
IF:
6
论文数:
5.2K
被引数:
1.1W

机构

H
huawei technologies
学者数:
3.3K
论文数: 2.9K
被引数: 1
F
fudan university
学者数:
11.8W
论文数: 7.7W
被引数: 121
C
Chinese University of Hong Kong
学者数:
3.4W
论文数: 3.2W
被引数: 5.6W
学者 查看更多机构
引用论文

引用论文

Studying brain activity during word-by-word interactions using wireless EEG
err2020-03-24
err0
errOAAI
errTatiana Goregliad Fjaellingsdal; Diana Schwenke; Esther Ruigendijk; Stefan Scherbaum; Martin Georg Bleichner
err分享
err收藏
Robust and Communication-Efficient Federated Learning From Non-i.i.d. Data
err2020-09-01
err1.0K
errOAAI
errSattler, Felix; Wiedemann, Simon; Mueller, Klaus-Robert; Samek, Wojciech
err分享
err收藏
err分享
err收藏
err分享
err收藏
Cyclopolyphosphines as ligands
err1966-08-01
err0
PREAI
errAlison Forster; C.S. Cundy; M. Green; F.G.A. Stone
err分享
err收藏
Synthesis and physical properties of new high temperature seignettomagnetics in the systems with perovskite type structure
err1997-12-01
err0
PREAI
errV. V. Gagulin; S. K. Korchagina; Yu. A. Shevchuk; N. V. Fadeeva; V. V. Bogatko
err分享
err收藏
A Machine Learning Approach for Blockchain-Based Smart Home Networks Security
err2021-05-01
err64
errOAAI
errKhan, Muhammad Adnan; Abbas, Sagheer; Rehman, Abdur; Saeed, Yousaf; Zeb, Asim; Uddin, M. Irfan; Nasser, Nidal; Ali, Asmaa
err分享
err收藏
err分享
err收藏
学者 查看更多内容