Return
Straggler-Aware In-Network Aggregation for Accelerating Distributed Deep Learning
DOI:10.1109/TSC.2023.3309318.png)
Abstract
En 中文
In-network aggregation facilitates accelerated distributed deep learning by utilizing a programmable switch to aggregate gradient packets. However, a straggler problem should be addressed to avoid performance degradation in terms of training time. In this paper, we propose a straggler-aware in-network aggregation (SAINA) scheme to mitigate the straggler problem while preventing accuracy degradation. In SAINA, the programmable switch aggregates local gradients of the fastest k workers to exclude stragglers and changes k adaptively to balance the tradeoff between training speed and accuracy. To this end, we design a switch-friendly convergence detection (SFCD) algorithm which detects a convergence point and determines k at the convergence point. SAINA is implemented over a software programmable switch and experimental results show that the accuracy of SAINA can reach a target accuracy up to 2.84x faster than the existing in-network aggregation scheme.
Keywords:
Distributed deep learning
in-network aggregation
programmable switch
straggler problem
Journal
IF:
5.8
Papers:
2.1K
Citations:
6.5K

