arrow
返回

Bridging the Gap Between Memory and Communication Efficiency on Distributed Deep Learning Systems

delete2021-01-01
delete0
delete
OA
AI
Z
Zhao, Shaofeng
刘博 封面图
刘博 (Bo Liu)
王
王芳 (Fang Wang) *
D
Dan Feng
DOI:10.1109/ACCESS.2021.3071579delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Large-scale distributed deep learning is of great importance in various applications. For data-parallel distributed training systems, limited hardware resources (e.g., GPU memory and interconnection bandwidth) often become a performance bottleneck, and it is necessary to consider the full utilization of multiple resources simultaneously, especially for extreme-scale deep neural networks. Although two different types of strategies, based on memory management and sparse communication, have been proposed to reduce the usage of resources, a naive combination of these two optimizations is impractical, since they cannot successfully coexist with each other. We therefore consider the idea of collaborative optimization in terms of both system memory and bandwidth resources, and propose a layer-centric memory-efficient distributed sparse communication mechanism called LaySA. Firstly, to tackle the memory ballooning issue caused by sparse communication, the existing memory reuse strategy is refined, and the data object of the memory optimization is augmented and redefined. Secondly, a mirror weight update mechanism is proposed to address the contradiction between memory management and sparse communication optimization for weight gradients. Our scheme, which involves the deep integration and collaborative execution of these two types of strategies, can fill the gap in relation to multiple resource optimization in distributed GPU-based training systems. Our experimental results show that the proposed collaborative optimization can significantly alleviate the memory pressure on the computing nodes, and improve both the resource utilization and the throughput of distributed training systems. Compared with baseline systems using only a single strategy, LaySA can help to reduce the system memory usage by up to 80.5%, and the overall training time of the neural network models on a single GPU is reduced by about 12.25%. Furthermore, LaySA can scale up the batch size of the datasets by an extremely large factor during distributed training, and the overall throughput is increased by more than 150%, meaning that our approach outperforms current systems that use memory or communication optimization mechanisms alone.
Keyword:
Training
Optimization
Memory management
Graphics processing units
Distributed databases
Bandwidth
Computational modeling
Deep learning
distributed training
intermediate data
memory management
sparse communication optimization
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

H
henan university of economics & law
学者数:
360
论文数: 403
被引数: 0
引用论文

引用论文

YAC transgene-mediated olfactory receptor gene choiceYAC转基因介导的嗅觉受体基因选择
err2000-02-01
err0
errOAAI
errFarah A.W. Ebrahimi; James Edmondson; Rodney Rothstein; Andrew Chess
err分享
err收藏
err分享
err收藏
err分享
err收藏
High-throughput machining using high average power ultrashort pulse lasers and ultrafast polygon scanner
err2016-03-04
err0
PREAI
errJoerg Schille; Lutz Schneider; André Streek; Sascha Kloetzer; Udo Loeschner
err分享
err收藏
没有更多内容