arrow
Return

Evaluation of Distributed Asynchronous Checkpointing in High-Performance Computing

delete2026-01-01
delete0
PRE
AI
R
Riccardo Scheda *
D
Domitilla Brandoni
L
Laura Cavalli
L
L. Morselli
DOI:10.1007/978-3-032-07612-0_13delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
During the training of large models, traditional checkpointing introduces significant overhead, as it requires pausing training while model states are copied from GPU memory to storage. At the same time, frequent checkpoint is needed to ensure an efficient use of resources. Efficient checkpointing is crucial for large-scale training of Artificial Intelligence (AI) models, especially on high-performance computing (HPC) systems. In this work, we evaluate distributed asynchronous checkpointing (DACP) applied on various LLM models on the Leonardo supercomputer, hosted by CINECA. By integrating asynchronous checkpointing, we enable overlapping data transfer operations with training iterations, significantly reducing training time. Our experiments span multiple LLM configurations, leveraging PyTorch DACP to optimize checkpointing frequency and minimize graphics processing unit (GPU) idle time. We evaluate this approach at scales of up to 256 GPUs using different model sizes. Results demonstrate a substantial reduction in checkpoint overhead, achieving up to a 6x improvement compared to synchronous methods. Our evaluations highlight the benefits of asynchronous checkpointing for large-scale training and provide insights into its practical deployment on HPC infrastructures.
Keywords:
Asynchronous checkpointing
Training
HPC
LLMs

Journal

H
HIGH PERFORMANCE COMPUTING WORKSHOPS, ISC HIGH PERFORMANCE 2025 INTERNATIONAL WORKSHOPS
IF:
0
Papers:
54
Citations:
0

Organization

C
cineca, italy
Scholars:
272
Papers: 260
Citations: 0