arrow
Return

Maximizing the Computation-Communication Overlap for Distributed Deep Learning With Approximate AllReduce

delete2026-03-27
delete0
PRE
AI
S
Shouxi Luo *
G
Gaolin Tang
X
Xue Liu
H
Huanlai Xing
DOI:10.1016/j.future.2026.108498delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
As is known, data-parallel distributed deep learning (DDL) calls for computation-communication overlap (CCO) aware communication optimization. The recent work of AQGB (Adaptive Quantized Gradient Broadcast) not only proposes the metric of ROW, i.e., the Ratio of Overlap time to Wait time, to quantify the optimization opportunity, but also designs an CCO-aware adaptive quantized gradient synchronization scheme for this purpose. Despite being efficient, AQGB is tailored for the case where training workers synchronize their gradients via direct broadcasting, thus unable to support DDL workloads relying on Ring-AllReduce and Halving Doubling (HD)-AllReduce based gradient synchronization.
Keywords:
Computation-Communication Overlap
Distributed Deep Learning
Gradient Synchronization
Approximate AllReduce
Adaptive Quantization

Journal

F
Future Generation Computer Systems
IF:
0
Papers:
642
Citations:
0

Organization

S
southwest jiaotong university
Scholars:
9.2K
Papers: 3.2K
Citations: 0