Return
Approximate Gradient Synchronization With Adaptive Quantized Gradient Broadcast
DOI:10.1016/j.future.2025.107983.png)
Abstract
En 中文
Given the layered training workflow of deep neural networks, recent advances have shown that, by splitting gradients into blocks and rearranging their transmission, distributed deep learning (DDL) workers can overlap parts of the communication with computation to hide the overhead of model synchronization. However, not all communication can be masked perfectly (e.g., that of the first-layer gradients). A promising solution is to transmit quantized gradients instead of raw values to eliminate the communication bottleneck further. In this paper, we propose AQGB, Adaptive Quantized Gradient Broadcast, to accelerate the convergence of data-parallel distributed training through designing efficient multi-level quantization and flexible quantization ratio control. Distinguished from existing fixed-quantization schemes, AQGB can adjust the level of quantization respecting the network state and the training progress to maximize the computation-communication overlap (CCO), which is quantified by a novel metric ROW (the Ratio of Overlap time to Wait time). Compared with no-quantization and 4bit-fixed QSGD quantization, AQGB could accelerate the convergence speed of training (regarding the time to converge) by about 3.15 × and 1.24 ×, respectively.
Journal
F
IF:
0
Papers:
642
Citations:
0
Organization
No organization information available

