Return
SIM: Accelerating Distributed DNN Training by Exploring Gradient Similarity
DOI:10.1109/TON.2025.3633713.png)
Abstract
En 中文
Synchronous stochastic gradient descent (SSGD) has been widely used in distributed deep learning. However, since the local gradients need to be shared among workers at every iteration, SSGD performance is significantly influenced by network bottlenecks caused by either heterogeneous environment or bandwidth contention. To solve this problem, asynchronous parallel (ASP) strategy allows each worker to update parameters independently without synchronization, while suffering from accuracy loss and convergence inefficiency. In this paper, we propose a novel similarity-based synchronization scheme called SIM, which mitigates the impact of network bottlenecks and ensures convergence efficiency. Specifically, SIM reduces the number of aggregation workers based on the gradient similarity between global and local gradients, therefore shrinking the waiting time for the stragglers. We provide a theoretical analysis of convergence efficiency and conduct large-scale testbed experiments on CIFAR-10 and SQUAD dataset. The experimental results show that SIM reduces the convergence time of four classical deep learning models by up to 40%.
Keywords:
Training
Convergence
Synchronization
Computational modeling
Deep learning
Accuracy
Artificial neural networks
Artificial intelligence
Stochastic processes
Servers
Distributed deep learning
straggler
synchronization scheme
gradient similarity
Journal
I
IF:
0
Papers:
543
Citations:
0

