Return
Priority-Driven Flow Scheduling for Distributed Large Language Model Training
DOI:10.23919/cje.2025.00.349.png)
Abstract
En 中文
Large language model (LLM) training is highly reliant on efficient communication coordination among distributed accelerators. While existing approaches focus on independently optimizing specific parallelism strategies, they lack systematic prioritization across different communication patterns, leading to suboptimal training performance. In this paper, we propose a hybrid parallelism priority assignment framework named HyPA. HyPA employs offline surrogate modeling for closed-form priority optimization and online parameter sensing for dynamic environmental adaptation, which enables adaptive bandwidth allocation and congestion control without modifying network infrastructure. Through a comprehensive evaluation on realistic training workloads, HyPA achieves significant improvements in the job completion time for both dense and sparse LLM models. There is a reduction in the job completion time of up to 18.59% in micro-benchmarks and up to 16% in large-scale training deployments.
Keywords:
Modeling
Training
Parallel processing
Large language models
Fluid flow
Optimization
Schedules
Scheduling
Calculators
Servers
Distributed training
Flow scheduling
Pipeline parallelism
Data parallelism
Communication optimization
Journal
C
IF:
3
Papers:
97
Citations:
1.7K
Organization
Cited Papers
No cited papers available

