arrow
Return

Communication Optimization for Large Language Model Distributed Training: A Systematic Survey

delete2026-01-01
delete0
PRE
AI
S
Sheng Li
J
Jun Zhu
Y
Yuanhao He *
H
Hanguang Luo
T
Tao Zou
G
Geyang Xiao
D
Donghui Lu
DOI:10.1007/978-981-95-3324-4_26delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The emergence of trillion-parameter large language models(LLMs) like GPT-4 and PaLM-2 has fundamentally transformed artificial general intelligence capabilities. However, their inefficient synchronization strategy and suboptimal network resource utilization together limit the system's scalability and training efficiency. This survey provides a systematic categorization of communication optimization techniques. It introduces a hierarchical framework that differentiates between logical strategy optimization and physical resource orchestration. Strategy optimizations focus on adaptive synchronization and gradient compression to reduce communication volume. Resource optimizations improve task scheduling, network topology, and hardware adaptation to boost physical efficiency. This work classifies current optimization strategies for large-scale distributed training systems. Further, it discusses potential research directions for improving communication efficiency in next-generation LLM distributed training systems.
Keywords:
communication optimization
distributed training
distributed parallelism
large scale
Large Language Model

Journal

P
PROCEEDINGS OF THE 15TH INTERNATIONAL CONFERENCE ON COMPUTER ENGINEERING AND NETWORKS, VOL II
IF:
0
Papers:
44
Citations:
0

Organization

Z
zhejiang laboratory
Scholars:
330
Papers: 195
Citations: 49
S
shenzhen university
Scholars:
4.5W
Papers: 3.4W
Citations: 72