Return
Maximizing the Computation-Communication Overlap for Distributed Deep Learning With Approximate AllReduce
DOI:10.1016/j.future.2026.108498.png)
Abstract
En 中文
As is known, data-parallel distributed deep learning (DDL) calls for computation-communication overlap (CCO) aware communication optimization. The recent work of AQGB (Adaptive Quantized Gradient Broadcast) not only proposes the metric of ROW, i.e., the Ratio of Overlap time to Wait time, to quantify the optimization opportunity, but also designs an CCO-aware adaptive quantized gradient synchronization scheme for this purpose. Despite being efficient, AQGB is tailored for the case where training workers synchronize their gradients via direct broadcasting, thus unable to support DDL workloads relying on Ring-AllReduce and Halving Doubling (HD)-AllReduce based gradient synchronization.
Keywords:
Computation-Communication Overlap
Distributed Deep Learning
Gradient Synchronization
Approximate AllReduce
Adaptive Quantization
Journal
F
IF:
0
Papers:
642
Citations:
0

