arrow
Return

Execution-Delay-Balanced Pipeline-Parallelism-Based Distributed Model Training for Artificial Intelligence Data Centers Interconnected by Optical Networks

delete2026-06-17
delete0
PRE
AI
J
Jingjie Xin
李
李昕 (Xin Li)
D
Daniel C. Kilper
S
Shanguo Huang
DOI:10.1109/jiot.2026.3704665delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Using pipeline parallelism to achieve distributed model training across artificial intelligence data centers (AIDCs) is feasible, where a Transformer-based distributed model training (TDMT) task is partitioned into multiple sequential stages and offloaded to suitable AIDCs. However, due to the uncertainty in the number and locations of partitioning points, determining suitable partitioning and offloading decisions to complete the training within the delay requirements is challenging. This article proposes an execution-delay-balanced pipeline-parallelism-based distributed model training (EP-DMT) scheme, which jointly schedules computing and network resources to determine partitioning and offloading decisions for TDMT tasks. The goal is to minimize the bubble time, thereby improving training efficiency. A heuristic algorithm is developed to determine each stage by estimating the basic number of stages and the execution delay of stages based on resource states. The selection of AIDCs and routing and wavelength allocation (RWA) solutions for stages is also optimized. Simulation results show that EP-DMT achieves a higher acceptance ratio compared with the benchmarks and a low bubble time.
Keywords:
Artificial intelligence data centers (AIDCs)
distributed model training (DMT)
optical networks
pipeline parallelism
task offloading

Journal

IEEE Internet of Things Journal cover
IEEE Internet of Things Journal
IF:
8.9
Papers:
1.4W
Citations:
7.8W

Organization

B
beijing university of posts and telecommunications
Scholars:
2.3K
Papers: 861
Citations: 0
T
Trinity College Dublin
Scholars:
2.4W
Papers: 1.9W
Citations: 2.7W
Cited Papers

Cited Papers

A Survey on Efficient Training of Transformers
err2023-08-01
err0
errOAAI
errBohan Zhuang; Jing Liu; Zizheng Pan; Haoyu He; Yuetian Weng; Chunhua Shen
errShare
errSave
Merak: An Efficient Distributed DNN Training Framework With Automated 3D Parallelism for Giant Foundation Models
err2023-05-01
err18
errOAAI
errLai, Zhiquan; Li, Shengwei; Tang, Xudong; Ge, Keshi; Liu, Weijie; Duan, Yabo; Qiao, Linbo; Li, Dongsheng
errShare
errSave
Distributed Model Training Based on Data Parallelism in Edge Computing-Enabled Elastic Optical Networks
err2021-04-01
err17
PREAI
errLi, Yajie; Zeng, Zebin; Li, Jun; Yan, Boyuan; Zhao, Yongli; Zhang, Jie
errShare
errSave
errShare
errSave
Nvidia Hopper GPU: Scaling Performance
err
IF0
err2022-08-21
err0
PREAI
errJack Choquette
errShare
errSave
Optimizing Multi-GPU Parallelization Strategies for Deep Learning Training
err2019-09-01
err0
errOAAI
errSaptadeep Pal; Eiman Ebrahimi; Arslan Zulfiqar; Yaosheng Fu; Victor Zhang; Szymon Migacz; David Nellans; Puneet Gupta
errShare
errSave
errShare
errSave
err
IF0
err
err0
PREAI
err
errShare
errSave
PyTorch distributed
err2020-09-14
err0
PREAI
errShen Li; Yanli Zhao; Rohan Varma; Omkar Salpekar; Pieter Noordhuis; Teng Li; Adam Paszke; Jeff Smith; Brian Vaughan; Pritam Damania; Soumith Chintala
errShare
errSave
researcher View more