Return
AIRP: Accelerating Multi-Tenant Distributed Learning With In-Network Resource Pooling
DOI:10.1109/TON.2026.3676228.png)
Abstract
En 中文
The increasing popularity of large models and datasets has highlighted the significance of distributed training networks. As gradient synchronization generates substantial traffic, in-network aggregation (INA) has emerged as a solution to offload aggregation onto the switch, alleviating network congestion and accelerating distributed training. However, the limited memory capacity of the INA switch becomes a potential bottleneck as computation shifts into the network, especially in multi-tenant scenarios. To address this bottleneck and enhance network throughput, we propose the Aggregation with In-network Resource Pooling (AIRP) framework. Unlike existing approaches that optimize individual switches in a localized manner, AIRP takes a holistic view and efficiently pools switch memory resources across the entire network, allocating them to multiple tenants. Evaluation using the ns-3 simulator and P4 testbed demonstrates that AIRP can accelerate the training of various models, including computer vision and language models. The experimental results show that AIRP outperforms existing INA approaches by up to 7 times in terms of network throughput in multi-tenant scenarios, while also achieving great flexibility and efficiency in deployment.
Keywords:
Switches
Memory management
Training
Bandwidth
Throughput
Resource management
Servers
Random access memory
Atmospheric modeling
Computational modeling
Distributed learning
in-network aggregation
resource pooling
maximum flow
Journal
I
IF:
0
Papers:
543
Citations:
0

