arrow
Return

Priority-Driven Flow Scheduling for Distributed Large Language Model Training

delete2026-05-01
delete0
PRE
AI
T
Tianshi Wang
Y
Yiran Zhang *
Z
Zhang, Yiao
Z
Zhang, Qiyang
Z
Zhou, Ao
S
Shangguang Wang
DOI:10.23919/cje.2025.00.349delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Large language model (LLM) training is highly reliant on efficient communication coordination among distributed accelerators. While existing approaches focus on independently optimizing specific parallelism strategies, they lack systematic prioritization across different communication patterns, leading to suboptimal training performance. In this paper, we propose a hybrid parallelism priority assignment framework named HyPA. HyPA employs offline surrogate modeling for closed-form priority optimization and online parameter sensing for dynamic environmental adaptation, which enables adaptive bandwidth allocation and congestion control without modifying network infrastructure. Through a comprehensive evaluation on realistic training workloads, HyPA achieves significant improvements in the job completion time for both dense and sparse LLM models. There is a reduction in the job completion time of up to 18.59% in micro-benchmarks and up to 16% in large-scale training deployments.
Keywords:
Modeling
Training
Parallel processing
Large language models
Fluid flow
Optimization
Schedules
Scheduling
Calculators
Servers
Distributed training
Flow scheduling
Pipeline parallelism
Data parallelism
Communication optimization

Journal

C
Chinese Journal of Electronics
IF:
3
Papers:
97
Citations:
1.7K

Organization

B
beijing university of posts & telecommunications
Scholars:
1.4W
Papers: 1.2W
Citations: 9
P
peking university
Scholars:
11.9W
Papers: 8.7W
Citations: 146
Cited Papers

Cited Papers

No cited papers available