arrow
Return

NetPlacer+: Model Parallelism Based on Load Balance in Distributed Deep Learning

delete2025-01-01
delete0
PRE
AI
Y
Yunqi Gao
Z
Zechao Zhang
B
Bing Hu
M
Mahdi Boloursaz Mashhadi
A
A-Long Jin
P
Pei Xiao
DOI:10.1109/TETCI.2025.3543765delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The importance of Model Parallelism in Distributed Deep Learning continues to grow due to the increase in the Deep Neural Network (DNN) scale and the demand for higher training speed. Different from all the existing works, we propose a model-parallel strategy called NetPlacer+ based on load balance. The major idea in NetPlacer+ is to partition the DNN model into multiple devices by balancing each device's computation and communication load. We build the mathematical model of NetPlacer+. We transform the mathematical model of NetPlacer+ and obtain its approximate optimal solution using the interior point method. Extensive experiments in two GPU clusters and eight modern DNNs are conducted to verify the effectiveness of NetPlacer+. Experimental results show that the model-parallel strategy of NetPlacer+ achieves up to 1.25x speedup compared to NVIDIA's DLPlacer.
Keywords:
Distributed deep learning
model parallelism
load balance

Journal

I
IEEE Transactions on Emerging Topics in Computational Intelligence
IF:
6.5
Papers:
1.4K
Citations:
4.5K

Organization

U
University of Surrey
Scholars:
1.2W
Papers: 1.3W
Citations: 22
Z
zhejiang university
Scholars:
17.6W
Papers: 12.1W
Citations: 152