arrow
返回

APapo: An asynchronous parallel optimization method for DNN models

delete2024-03-01
delete1
PRE
AI
S
Shuai Liu
T
Tao Ju *
DOI:10.1016/j.future.2023.11.004delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
To address the challenges related to segmentation complexity, high memory usage, extended training duration, and low equipment utilization in parallel optimization of large-scale deep neural network (DNN) models, this paper proposes an asynchronous parallel optimization method APapo. Firstly, a multi-iteration asynchronous pipeline parallel scheduling was established for model parallel computing tasks, controlling the specific scheduling process of micro-batch units to address gradient delay updating during asynchronous iteration. Secondly, combined with the given network model and hardware configuration, a dynamic programming strategy for computing resources and model tasks was designed to achieve dynamic segmentation of model computing tasks and optimal matching of computing resources. Finally, an optimization strategy for runtime scheduling of computing resources and model tasks was developed, using improved device streams to maximize the overlap between computing and communication, thus improving the utilization rate of computing resources and reducing training time. Experimental results show that the APapo method achieves fine-grained task segmentation, maximizes the utilization rate of each GPU computing resource, and on average improves the training speed of large-scale deep neural network models by 2.8 times while maintaining the training accuracy of the model compared to existing parallel optimization methods.
Keyword:
DNN model parallelism
Model segmentation
Asynchronous pipeline parallelism
Augmented antichain
Computation-communication overlap

期刊

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
论文数:
6.8K
被引数:
2.3W

机构

L
Lanzhou Jiaotong University
学者数:
6.3K
论文数: 3.6K
被引数: 4.2K
引用论文

引用论文

vPipe: A Virtualized Acceleration System for Achieving Efficient and Scalable Pipeline Parallel DNN Training
err2022-03-01
err22
errOAAI
errZhao, Shixiong; Li, Fanxin; Chen, Xusheng; Guan, Xiuxian; Jiang, Jianyu; Huang, Dong; Qing, Yuhao; Wang, Sen; Wang, Peng; Zhang, Gong; Li, Cheng; Luo, Ping; Cui, Heming
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
Transplantation of CD34+ peripheral blood progenitor cells after high- dose chemotherapy for patients with advanced multiple myeloma
err1995-07-01
err0
errOAAI
errG Schiller; R Vescio; C Freytes; G Spitzer; F Sahebi; M Lee; CH Wu; J Cao; JC Lee; CH Hong
err分享
err收藏
err分享
err收藏
学者 查看更多内容