Return
Accelerating Distributed Training Through In-Network Aggregation and Route Selection
DOI:10.1109/tc.2026.3709852.png)
Abstract
En 中文
Deep learning has brought about a revolutionary transformation in network applications, particularly in domains like e-commerce and online advertising. Distributed training (DT), as a critical means to expedite model training, has progressively emerged as a key foundational infrastructure for such applications. However, with the rapid advancement of hardware accelerators, the bottleneck in DT has shifted from computation to communication. To this end, one promising approach called in-network aggregation (INA) is proposed to alleviate the bandwidth bottleneck by aggregating gradients using programmable hardware (<i>e</i>.<i>g</i>., Intel Tofino switches). Regrettably, current INA works primarily focus on single-PS architecture and do not fully address the communication bottleneck caused by limited PS ingress bandwidth. To bridge this gap, we propose InArt, the first work introducing INA with route selection in a multi-PS architecture. To accommodate traffic dynamics, InArt adopts a two-phase approach: splitting the DT task among multiple parameter servers and selecting appropriate routing schemes to harness INA capabilities fully. We propose Lagrange multiplier and randomized rounding algorithms for these phases, respectively. We implement InArt and evaluate its performance through experiments on physical platforms (Tofino switches) and Mininet emulation (P4 Software Switches). Compared with state-of-the-art solutions, experimental results show that InArt can reduce communication time by <inline-formula><tex-math notation="LaTeX">$\mathbf{66\%}$</tex-math></inline-formula> on average.
Keywords:
Distributed training
in-network aggregation
route selection
programmable switches
Journal
IF:
3.8
Papers:
5.3K
Citations:
9.8K

