Return
Execution-Delay-Balanced Pipeline-Parallelism-Based Distributed Model Training for Artificial Intelligence Data Centers Interconnected by Optical Networks
J
李
D
S
DOI:10.1109/jiot.2026.3704665.png)
Abstract
En 中文
Using pipeline parallelism to achieve distributed model training across artificial intelligence data centers (AIDCs) is feasible, where a Transformer-based distributed model training (TDMT) task is partitioned into multiple sequential stages and offloaded to suitable AIDCs. However, due to the uncertainty in the number and locations of partitioning points, determining suitable partitioning and offloading decisions to complete the training within the delay requirements is challenging. This article proposes an execution-delay-balanced pipeline-parallelism-based distributed model training (EP-DMT) scheme, which jointly schedules computing and network resources to determine partitioning and offloading decisions for TDMT tasks. The goal is to minimize the bubble time, thereby improving training efficiency. A heuristic algorithm is developed to determine each stage by estimating the basic number of stages and the execution delay of stages based on resource states. The selection of AIDCs and routing and wavelength allocation (RWA) solutions for stages is also optimized. Simulation results show that EP-DMT achieves a higher acceptance ratio compared with the benchmarks and a low bubble time.
Keywords:
Artificial intelligence data centers (AIDCs)
distributed model training (DMT)
optical networks
pipeline parallelism
task offloading
Journal
IF:
8.9
Papers:
1.4W
Citations:
7.8W
