arrow
Return

The Asynchronous Training Algorithm Based on Sampling and Mean Fusion for Distributed RNN

delete2020-01-01
delete1
delete
OA
AI
D
Dejiao Niu
T
Tianquan Liu
T
Tao Cai *
S
Shijie Zhou
DOI:10.1109/ACCESS.2019.2939851delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Training of large scale deep neural networks with distributed implementations is an effective way to improve the efficiency. However, high network communication cost for synchronizing gradients and parameters is a major bottleneck in distributed training. In this work, we propose an asynchronous training algorithm based on sampling and mean fusion for distributed recurrent neural network (RNN). In distributed RNN, multiple distributed neuron nodes and an interaction node work together to implement the training. The synchronization overhead is reduced by a unique asynchronous sampling strategy amongst the distributed neuron nodes. Then, in order to make up for the accuracy loss caused by the asynchronous parameter update, a mean fusion algorithm is proposed, where the interaction node averages all local parameters from the distributed neurons. We mathematically prove the convergence of the proposed algorithm. Experimental verification is performed on two language modeling benchmark datasets. The results demonstrate significant speed gains for distributed RNN, while the accuracy loss is less than 1x0025; on average.
Keywords:
Training
Neurons
Parallel processing
Synchronization
Computational modeling
Data models
Recurrent neural networks
Asynchronous training
distributed recurrent neural network
mean fusion
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

J
Jiangsu University
Scholars:
4.0W
Papers: 2.8W
Citations: 5.5W