arrow
Return

When In-Network Computing Meets Distributed Machine Learning

delete2024-09-01
delete0
PRE
AI
H
Haowen Zhu
姜文超 (Wenchao Jiang)
Q
Q.H. Hong
郭泽华 (Zehua Guo) *
DOI:10.1109/MNET.2024.3368138delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Emerging In-Network Computing (INC) technique provides a new opportunity to improve application's performance by using network programmability, computational capability, and storage capacity enabled by programmable switches. One typical application is Distributed Machine Learning (DML), which accelerates machine learning training by employing multiple works to train model parallelly. This paper introduces INC-based DML systems, analyzes performance improvement from using INC, and overviews current studies of INC-based DML systems. We also propose potential research directions for applying INC to DML systems.
Keywords:
Training
Computational modeling
Servers
Data models
Machine learning
Synchronization
Performance evaluation
Distributed Machine Learning
In-Network Computing
Machine Learning
Programmable Switch

Journal

IEEE Network cover
IEEE Network
IF:
6.3
Papers:
2.6K
Citations:
1.1W

Organization

S
singapore university of technology & design
Scholars:
2.8K
Papers: 3.6K
Citations: 5
B
beijing institute of technology
Scholars:
5.4W
Papers: 3.9W
Citations: 63