arrow
返回

Optimizing Multi-Level Checkpointing for Distributed Deep Learning Workloads on Cloud Spot VM Clusters

delete2024-01-01
delete0
delete
OA
AI
Y
Yonghyeon Cho
Y
Yoochan Kim *
K
Kihyun Kim
J
Jinwoo Kim
H
Hong-Yeon Kim
Y
Youngjae Kim *
DOI:10.1109/ACCESS.2024.3446770delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Spot Virtual Machines (Spot VMs) offer access to underutilized computing resources at significant discounts, sometimes up to 90% off regular on-demand pricing. For budget-conscious organizations, using clusters of Spot VMs is an effective strategy for training large-scale distributed deep learning (DDL) models. However, the risk of preemption by cloud providers poses a challenge, as it can result in the loss of unsaved data in memory and local storage. To mitigate this risk, one solution involves using networked storage systems for checkpoints, though their low write throughput can slow down training. An alternative approach is to use the memory of a remote, on-demand computing node for temporary checkpoint storage, balancing data protection with training efficiency. In this paper, we propose a novel approach, ACUTE, to optimize temporary checkpointing in the memory of on-demand nodes during DDL training. ACUTE includes three key optimizations: 1) Check-Mem, which reduces memory copying overhead on the training node; 2) Check-Trans, which accelerates checkpoint data transfer through parallel processing; and 3) Check-Pack, which eliminates unnecessary data unpacking and repacking. Implemented using PyTorch's distributed data-parallel library, ACUTE was evaluated against two other checkpointing schemes on AWS VM instances. Results show that ACUTE reduces makespan delay to nearly zero and achieves, on average, 43.30% faster checkpointing compared to a baseline multi-level checkpointing scheme, without compromising the precision of Deep Neural Network (DNN) models.
Keyword:
Checkpointing
Artificial neural networks
Parallel processing
Deep learning
Cloud computing
Graphics processing units
Fault tolerance
Distributed deep learning
cloud computing
fault-tolerant systems
checkpoint and restart

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

L
lg corporation
学者数:
377
论文数: 221
被引数: 0
L
LG Electronics
学者数:
781
论文数: 635
被引数: 0
S
Sogang University
学者数:
4.6K
论文数: 4.4K
被引数: 4.0K
学者 查看更多机构
引用论文

引用论文

A Comprehensive Study of Load Balancing Approaches in the Cloud Computing Environment and a Novel Fault Tolerance Approach
err2020-01-01
err32
errOAAI
errShahid, Muhammad Asim; Islam, Noman; Alam, Muhammad Mansoor; Su'ud, Mazliham Mohd; Musa, Shahrulniza
err分享
err收藏
RANKL expression is related to treatment outcome of patients with localized, high‐grade osteosarcoma
err2010-12-17
err0
PREAI
errJun Ah Lee; Jun Soo Jung; Dong Ho Kim; Jung Sub Lim; Min Suk Kim; Chang‐Bae Kong; Won Seok Song; Wan Hyeong Cho; Dae‐Geun Jeon; Soo‐Yong Lee; Jae‐Soo Koh
err分享
err收藏
Floquet-Bloch operator for the Bose-Hubbard model with static field
err2003-11-25
err0
PREAI
errAndrey R. Kolovsky; Andreas Buchleitner
err分享
err收藏
err分享
err收藏
学者 查看更多内容