arrow
Return

swPredictor: A data-driven performance model for distributed data parallelism training on large-scale HPC clusters

delete2025-11-01
delete0
PRE
AI
X
Xianyu Zhu
R
Ruohan Wu
J
Junshi Chen *
安虹 (Hong An)
DOI:10.1016/j.peva.2025.102530delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Given the complexity of heterogeneous architectures and multi-node collaboration, largescale HPC (High-Performance Computing) clusters pose challenges in resource utilization and performance optimization during distributed data parallelism (DDP) training. Performance modeling aims to identify application bottlenecks and guide algorithm design, but existing performance models rarely consider the impact of system architecture on communication performance or provide a systematic analysis of distributed training. To address these issues, this paper proposes swPredictor, a data-driven performance model devised for accurately predicting the performance of DDP training. First, an original performance dataset is developed based on various communication patterns at runtime to avoid systematic errors. Subsequently, a novel multi-branch module FNO-Inception is proposed, combining FNO (Fourier Neural Operator) layer with Inception structure to simultaneously utilize various frequency features. Finally, by introducing the FNO-Inception module, a novel regression model FI-Net is constructed to fit complex nonlinear relationships. The experimental results demonstrate that FI-Net can accurately predict the performance of DDP training on the Sunway OceanLight supercomputer with an overall MAPE of 0.93%, which outperforms the other baseline models.
Keywords:
High-performance computing
Performance modeling
Deep learning
Distributed training

Journal

P
Performance Evaluation
IF:
0.8
Papers:
38
Citations:
851

Organization

C
chinese academy of sciences
Scholars:
56.5W
Papers: 44.9W
Citations: 704