arrow
返回

A Communication Efficient ADMM-based Distributed Algorithm Using Two-Dimensional Torus Grouping AllReduce

delete2023-01-02
delete3
delete
OA
AI
G
Guozheng Wang *
Y
Yongmei Lei
Z
Zeyu Zhang
C
Cunlu Peng
DOI:10.1007/s41019-022-00202-7delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Large-scale distributed training mainly consists of sub-model parallel training and parameter synchronization. With the expansion of training workers, the efficiency of parameter synchronization will be affected. To tackle this problem, we first propose 2D-TGA, a grouping AllReduce method based on the two-dimensional torus topology. This method synchronizes the model parameters by grouping and makes full use of bandwidth. Secondly, we propose a distributed algorithm, 2D-TGA-ADMM, which combines the 2D-TGA with the alternating direction method of multipliers (ADMM). It focuses on sub-model training and reduces the wait time among workers in the synchronization process. Finally, experimental results on the Tianhe-2 supercomputing platform show that compared with the MPI_Allreduce, the 2D-TGA could shorten the synchronization wait time by 33%.
Keyword:
ADMM
Grouping AllReduce
Two-dimensional torus topology
Synchronous algorithm

期刊

D
Data Science and Engineering
IF:
4.6
论文数:
249
被引数:
665

机构

S
shanghai university
学者数:
3.9W
论文数: 2.7W
被引数: 52