Return
MultiGS: Multi-Dimensional Information-Aware Gradient Sparsification
DOI:10.1109/TCSVT.2025.3636885.png)
Abstract
En 中文
Gradient sparsification (GS) is an effective method for reducing communication overhead in distributed training. For the first time, we introduce the concept of Multi-dimensional information into GS and propose a new gradient sparsification method named Multi-dimensional information-aware Gradient sparsification (MultiGS), which achieves high compression ratio with negligible accuracy loss and is applicable to mainstream network architectures. MultiGS reconstructs the layer-wise gradient by combining the high-frequency components of the local gradient and the low-frequency components of the sparsified global gradient that effectively addresses the issue of stale gradients and alleviates model bifurcation. Through the convergence proof of MultiGS for smooth non-convex problems and comparison with momentum SGD in convergence speed, we show that such new perspective approach is theoretically reasonable and practically effective. As validated with several mainstream model families (i.e., ResNets, VGGNet, LSTM, Vision Transformer, and Large Language Models), our MultiGS shows better accuracy over previous GS methods. Moreover, empirical results show that when a sufficient number of training nodes are available, MultiGS accelerates the distributed training by more than <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$3\times $ </tex-math></inline-formula>, which is better than existing sparsification method.
Keywords:
Distributed training
gradient sparsification
neural network acceleration
Journal
IF:
11.1
Papers:
612
Citations:
3.1W

