arrow
Return

Straggler-Aware In-Network Aggregation for Accelerating Distributed Deep Learning

delete2023-11-01
delete2
PRE
AI
H
Hochan Lee
J
Jaewook Lee
H
Heewon Kim
S
Sangheon Pack *
DOI:10.1109/TSC.2023.3309318delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In-network aggregation facilitates accelerated distributed deep learning by utilizing a programmable switch to aggregate gradient packets. However, a straggler problem should be addressed to avoid performance degradation in terms of training time. In this paper, we propose a straggler-aware in-network aggregation (SAINA) scheme to mitigate the straggler problem while preventing accuracy degradation. In SAINA, the programmable switch aggregates local gradients of the fastest k workers to exclude stragglers and changes k adaptively to balance the tradeoff between training speed and accuracy. To this end, we design a switch-friendly convergence detection (SFCD) algorithm which detects a convergence point and determines k at the convergence point. SAINA is implemented over a software programmable switch and experimental results show that the accuracy of SAINA can reach a target accuracy up to 2.84x faster than the existing in-network aggregation scheme.
Keywords:
Distributed deep learning
in-network aggregation
programmable switch
straggler problem

Journal

IEEE Transactions on Services Computing cover
IEEE Transactions on Services Computing
IF:
5.8
Papers:
2.1K
Citations:
6.5K

Organization

K
Korea University
Scholars:
3.6W
Papers: 3.8W
Citations: 4.4W
P
Pukyong National University
Scholars:
6.1K
Papers: 6.5K
Citations: 6.3K