arrow
Return

SIM: Accelerating Distributed DNN Training by Exploring Gradient Similarity

delete2025-11-01
delete0
PRE
AI
Y
Ye Jin
Y
Yijun Li
W
Wenliang Li
X
Xiaojuan Lu
Q
Qichen Su
黄家玮 (Jiawei Huang) *
王健鑫 (Jianxin Wang)
DOI:10.1109/TON.2025.3633713delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Synchronous stochastic gradient descent (SSGD) has been widely used in distributed deep learning. However, since the local gradients need to be shared among workers at every iteration, SSGD performance is significantly influenced by network bottlenecks caused by either heterogeneous environment or bandwidth contention. To solve this problem, asynchronous parallel (ASP) strategy allows each worker to update parameters independently without synchronization, while suffering from accuracy loss and convergence inefficiency. In this paper, we propose a novel similarity-based synchronization scheme called SIM, which mitigates the impact of network bottlenecks and ensures convergence efficiency. Specifically, SIM reduces the number of aggregation workers based on the gradient similarity between global and local gradients, therefore shrinking the waiting time for the stragglers. We provide a theoretical analysis of convergence efficiency and conduct large-scale testbed experiments on CIFAR-10 and SQUAD dataset. The experimental results show that SIM reduces the convergence time of four classical deep learning models by up to 40%.
Keywords:
Training
Convergence
Synchronization
Computational modeling
Deep learning
Accuracy
Artificial neural networks
Artificial intelligence
Stochastic processes
Servers
Distributed deep learning
straggler
synchronization scheme
gradient similarity

Journal

I
IEEE Transactions on Networking
IF:
0
Papers:
543
Citations:
0

Organization

G
Guangxi University
Scholars:
4.1K
Papers: 1.3K
Citations: 3.2W
C
central south university
Scholars:
1.9W
Papers: 5.7K
Citations: 3