arrow
Return

Dynamic layer-wise sparsification for distributed deep learning

delete2023-10-01
delete4
PRE
AI
张浩 (Hao Zhang)
T
Tingting Wu *
Z
Zhifeng Ma
F
Feng Li
刘洁 (Jie Liu)
DOI:10.1016/j.future.2023.04.022delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Distributed stochastic gradient descent (SGD) algorithms are becoming popular in speeding up deep learning model training by employing multiple computational devices (named workers) parallelly. Top-k sparsification, a mechanism where each worker only communicates a small number of largest gradients (by absolute value) and accumulates the rest locally, is one of the most basic and high-profile practices to reduce communication overhead. However, the theoretical implementation (Global Top-k SGD) ignoring the layer-wise structure of neural networks has low training efficiency, since the topk operation requiring the whole gradients impedes parallelism of computation and communication. The practical implementation (Layer-wise Top-k SGD) solves the parallelism problem, but hurts the performance of the trained model due to the deviation from the theoretically optimal solution. In this paper, we solve this contradiction by introducing a Dynamic Layer-wise Sparsification (DLS) mechanism and its extensions, DLS(s). DLS(s) efficiently adjusts the sparsity ratios of the layers to make the uploaded threshold of each layer automatically tend to be the unified global one, so as to retain the good performance of Global Top-k SGD and the high efficiency of Layer-wise Top-k SGD. The experimental results show that DLS(s) outperforms Layer-wise Top-k SGD in performance, and performs close to Global Top-k SGD yet have much less training time. (c) 2023 Elsevier B.V. All rights reserved.
Keywords:
Distributed deep learning
Parallel training
Stochastic gradient descent
Stochastic optimization
Gradient sparsification
Top-k

Journal

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
Papers:
6.8K
Citations:
2.3W

Organization

H
harbin institute of technology
Scholars:
8.0W
Papers: 6.6W
Citations: 66