Return
When In-Network Computing Meets Distributed Machine Learning
DOI:10.1109/MNET.2024.3368138.png)
Abstract
En 中文
Emerging In-Network Computing (INC) technique provides a new opportunity to improve application's performance by using network programmability, computational capability, and storage capacity enabled by programmable switches. One typical application is Distributed Machine Learning (DML), which accelerates machine learning training by employing multiple works to train model parallelly. This paper introduces INC-based DML systems, analyzes performance improvement from using INC, and overviews current studies of INC-based DML systems. We also propose potential research directions for applying INC to DML systems.
Keywords:
Training
Computational modeling
Servers
Data models
Machine learning
Synchronization
Performance evaluation
Distributed Machine Learning
In-Network Computing
Machine Learning
Programmable Switch
Journal
IF:
6.3
Papers:
2.6K
Citations:
1.1W

