Return
DINA: Toward Determined In-Network Aggregation for Distributed Machine Learning
DOI:10.1109/TON.2025.3535707.png)
Abstract
En 中文
Distributed Machine Learning (DML) utilizes parallel computation on multiple training nodes to accelerate machine learning model training. Parameter Server (PS) is a typical DML enabler and is widely used in industry and academia. Existing works propose to apply the emerging In-Network Aggregation (INA) technique to improve model training efficiency by offloading the whole gradient aggregation process in PS from hosts to programmable switches. However, existing INA systems may suffer from undetermined model training efficiency and service quality, given that many gradient aggregation processes are still performed by the server under irrational gradient aggregation strategies. In this paper, we propose a Deterministic In-Network Aggregation (DINA) scheme to improve model training efficiency by enhancing the efficiency of INA utilization in DML. Our key observation is to further increase worker sending rates by reducing gradient packets’ RTT (i.e., realizing packet sub-RTT). Based on this observation, DINA rationally selects the optimal global gradient aggregation switch depending on the switches’ available memory, worker sending rate, and server processing capacity. As a result, DINA reduces the dependence of INA systems on the server, improves worker sending rates, and mitigates network traffic load. We formulate the sub-RTT-INA-based gradient aggregation problem as a mixed-integer nonlinear programming problem. To efficiently solve the problem, we simplify it by transforming the nonlinear constraints into linear constraints and propose a mixed solution that combines randomized rounding and heuristic mechanisms. Simulation results show that DINA can provide determined training by reducing communication time by 12%-17% and network load by 28%-50% compared with existing solutions, thus taking full advantage of INA and realizing a determined INA service.
Keywords:
Distributed machine learning (DML)
in-network aggregation (INA)
sub-round-trip-time (RTT)
gradient routing
Journal
I
IF:
0
Papers:
543
Citations:
0

