arrow
Return

Parallel I/O analysis in distributed deep learning applications on high-performance computing

delete2025-11-05
delete0
delete
OA
AI
E
Edixon Párraga *
B
Betzabeth León
S
Sandra Méndez
D
Dolores Rexachs
E
Emilio Luque
DOI:10.1007/s11227-025-07986-1delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Distributed deep learning (DDL) applications generate heavy input/output (I/O) workloads that can create bottlenecks in high-performance computing (HPC) systems. Their optimal I/O configuration depends on factors such as access patterns, storage hardware, dataset size, and execution scale. This study proposes a systematic methodology for characterizing and optimizing I/O behavior in DDL applications, represented through the deep learning I/O benchmark (DLIO), and validated with the real DeepGalaxy application. We evaluate access modes, file formats, and Lustre file system configurations, demonstrating that stripe counts optimized for the access pattern and application scale can reduce I/O and execution times, achieving up to 18 GiB/s of bandwidth and a 5X increase in IOPS. HDF5 provides balanced performance, while TFRecord stands out in bandwidth-intensive scenarios. Shared access minimizes contention and improves scalability in multi-node executions. The results are consolidated into configuration guidelines that offer practical recommendations for practitioners to tune DDL applications for efficient execution in HPC environments.
Keywords:
Distributed deep learning
Parallel I/O
HPC cluster
I/O behavior patterns
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

T
The Journal of Supercomputing
IF:
0
Papers:
647
Citations:
0

Organization

C
computer sciences department
Scholars:
4
Papers: 5
Citations: 0