arrow
返回

MIPD: An Adaptive Gradient Sparsification Framework for Distributed DNNs Training

delete2022-01-01
delete8
delete
OA
AI
Z
Zhaorui Zhang *
DOI:10.1109/TPDS.2022.3154387delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Asynchronous training based on the parameter server architecture is widely used for scaling up the DNN training over large datasets and DNN models. Communication has been identified as the major bottleneck when deploying the DNN training over the large-scale distributed deep learning systems. Recent studies try to reduce the communication traffic through gradient sparsification and quantization approaches. We identify three limitations in previous studies. First, the fundamental guideline for gradient sparsification of their work is the magnitude of the gradient. However, the gradients' magnitude represents the current optimization direction while it cannot indicate the significance of the parameters, which potentially results in delayed updating for the significant parameters. Second, their gradient quantization methods based on the entire model often lead to error accumulation for gradients aggregation since the gradients from different layers of the DNN model follow different distributions. Third, previous quantization approaches are CPU intensive, which generates strong overhead for the server. We propose MIPD, an adaptive and layer-wise gradient sparsification framework that compresses the gradients based on model interpretability and probability distribution of gradients. MIPD compresses the gradients according to the corresponding significance of its parameters, which is defined by model interpretability. An Exponential Smoothing method is also proposed to compensate for the dropped gradients on the server to reduce the gradients error. MIPD proposes to update half of the parameters for each training step to reduce the CPU overhead of the server. It encodes the gradients based on their probability distribution, thereby minimizing the approximated errors. Extensive experimental results generated on the GPU cluster indicate that the proposed framework effectively improves the training performance of DNNs by up to 36.2%, which ensures high accuracy as compared to state-of-art solutions. Accordingly, the CPU and network usage of the server dropped by up to 42.0% and 32.7% respectively.
Keyword:
Training
Servers
Quantization (signal)
Convergence
Probability distribution
Degradation
Adaptation models
Exponential smoothing prediction
gradients sparsification
model interpretability
probability distribution
quantization

期刊

IEEE Transactions on Parallel and Distributed Systems 封面图
IEEE Transactions on Parallel and Distributed Systems
IF:
6
论文数:
5.2K
被引数:
1.1W

机构

U
University of Hong Kong
学者数:
4.1W
论文数: 3.9W
被引数: 10.1W
引用论文

引用论文

Knowledge Distillation: A Survey知识蒸馏: 一项调查
err2021-03-22
err1.5K
PREAI
errGou, Jianping; Yu, Baosheng; Maybank, Stephen J.; Tao, Dacheng
err分享
err收藏
Embedded importance watermarking for image verification in radiology
err2004-03-29
err0
errOAAI
errDomininc Osborne; D. Rogers; M. Sorell; Derek Abbott
err分享
err收藏
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容