返回
Bridging the Gap Between Memory and Communication Efficiency on Distributed Deep Learning Systems
DOI:10.1109/ACCESS.2021.3071579.png)
摘要
En 中文
Large-scale distributed deep learning is of great importance in various applications. For data-parallel distributed training systems, limited hardware resources (e.g., GPU memory and interconnection bandwidth) often become a performance bottleneck, and it is necessary to consider the full utilization of multiple resources simultaneously, especially for extreme-scale deep neural networks. Although two different types of strategies, based on memory management and sparse communication, have been proposed to reduce the usage of resources, a naive combination of these two optimizations is impractical, since they cannot successfully coexist with each other. We therefore consider the idea of collaborative optimization in terms of both system memory and bandwidth resources, and propose a layer-centric memory-efficient distributed sparse communication mechanism called LaySA. Firstly, to tackle the memory ballooning issue caused by sparse communication, the existing memory reuse strategy is refined, and the data object of the memory optimization is augmented and redefined. Secondly, a mirror weight update mechanism is proposed to address the contradiction between memory management and sparse communication optimization for weight gradients. Our scheme, which involves the deep integration and collaborative execution of these two types of strategies, can fill the gap in relation to multiple resource optimization in distributed GPU-based training systems. Our experimental results show that the proposed collaborative optimization can significantly alleviate the memory pressure on the computing nodes, and improve both the resource utilization and the throughput of distributed training systems. Compared with baseline systems using only a single strategy, LaySA can help to reduce the system memory usage by up to 80.5%, and the overall training time of the neural network models on a single GPU is reduced by about 12.25%. Furthermore, LaySA can scale up the batch size of the datasets by an extremely large factor during distributed training, and the overall throughput is increased by more than 150%, meaning that our approach outperforms current systems that use memory or communication optimization mechanisms alone.
Keyword:
Training
Optimization
Memory management
Graphics processing units
Distributed databases
Bandwidth
Computational modeling
Deep learning
distributed training
intermediate data
memory management
sparse communication optimization
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
Do Perceptions of Competence Mediate The Relationship Between Fundamental Motor Skill Proficiency and Physical Activity Levels of Children in Kindergarten?能力的感知是否可以介导幼儿园儿童的基本运动技能熟练程度与身体活动水平之间的关系?
Developing Wind and/or Solar Powered Crop Irrigation Systems for the Great Plains为大平原开发风能和/或太阳能作物灌溉系统
没有更多内容

