arrow
Return

Improving Automatic Parallel Training via Balanced Memory Workload Optimization

delete2024-08-01
delete0
delete
OA
AI
Y
Y X Wang *
Y
Youhe Jiang
X
Xupeng Miao
F
Fangcheng Fu
S
Shenhan Zhu
X
Xiaonan Nie
Y
Yaofeng Tu
崔斌 cover
崔斌 (Bin Cui)
DOI:10.1109/TKDE.2024.3370614delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Transformer models have emerged as the leading approach for achieving state-of-the-art performance across various application domains, serving as the foundation for advanced large-scale deep learning (DL) models. However, efficiently training these models across multiple GPUs remains a complex challenge due to the abundance of parallelism options. Existing DL systems either require manual efforts to design distributed training plans or limit parallelism combinations to a constrained search space. In this paper, we present Galvatron-BMW, a novel system framework that integrates multiple prevalent parallelism dimensions and automatically identifies the most efficient hybrid parallelism strategy. To effectively navigate this vast search space, we employ a decision tree approach for decomposition and pruning based on intuitive insights. We further utilize a dynamic programming search algorithm to derive the optimal plan. Moreover, to improve resource utilization and enhance system efficiency, we propose a bi-objective optimization workflow that focuses on workload balance. Our evaluations on different Transformer models demonstrate the capabilities of Galvatron-BMW in automating distributed training under varying GPU memory constraints. Across all tested scenarios, Galvatron-BMW consistently achieves superior system throughput, surpassing previous approaches that rely on limited parallelism strategies.
Keywords:
Parallel processing
Transformers
Computational modeling
Training
Data models
Memory management
Costs
distributed learning
automatic parallelism

Journal

IEEE Transactions on Knowledge and Data Engineering cover
IEEE Transactions on Knowledge and Data Engineering
IF:
10.4
Papers:
6.7K
Citations:
3.2W

Organization

C
Carnegie Mellon University
Scholars:
1.4W
Papers: 1.4W
Citations: 2.7W
P
peking university
Scholars:
11.7W
Papers: 8.7W
Citations: 146
Z
zte
Scholars:
419
Papers: 412
Citations: 0
researcher View more organizations