arrow
返回

Simplifying Distributed Neural Network Training on Massive Graphs: Randomized Partitions Improve Model Aggregation

delete2024-12-28
delete0
delete
OA
AI
J
Jiong Zhu *
A
Aishwarya Reganti
E
Edward W Huang
C
Charles Dickens
N
Nikhil Rao
K
Karthik Subbian
D
Danai Koutra
DOI:10.1145/3701563delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Distributed graph neural network (GNN) training facilitates learning on massive graphs that surpass the storage and computational capabilities of a single machine. Traditional distributed frameworks strive for performance parity with centralized training by maximally recovering cross-instance node dependencies, relying either on inter-instance communication or periodic fallback to centralized training. However, these processes create overhead and constrain the scalability of the framework. In this work, we propose a streamlined framework for distributed GNN training that eliminates these costly operations, yielding improved scalability, convergence speed, and performance over state-of-the-art approaches. Our framework (1) comprises independent trainers that asynchronously learn local models from locally available parts of the training graph and (2) synchronizes these local models only through periodic (time-based) model aggregation. Contrary to prevailing belief, our theoretical analysis shows that it is not essential to maximize the recovery of cross-instance node dependencies to achieve performance parity with centralized training. Instead, our framework leverages randomized assignment of nodes or super-nodes (i.e., collections of original nodes) to partition the training graph in order to enhance data uniformity and minimize discrepancies in gradient and loss function across instances. Experiments on social and e-commerce networks with up to 1.3 billion edges show that our proposed framework achieves state-of-the-art performance and 2.31x speedup compared to the fastest baseline despite using less training data.
Keyword:
graph neural networks
scalability
distributed learning
model aggregation training

期刊

ACM Transactions on Knowledge Discovery from Data 封面图
ACM Transactions on Knowledge Discovery from Data
IF:
4.8
论文数:
1.3K
被引数:
4.4K

机构

A
amazon.com
学者数:
699
论文数: 506
被引数: 8
University of California System 封面图
University of California System
学者数:
37.7W
论文数: 33.8W
被引数: 6.6K
U
University of Michigan
学者数:
6.4W
论文数: 5.3W
被引数: 124
学者 查看更多机构
引用论文

引用论文

Association of p21‐activated kinase‐1 activity with aggressive tumor behavior and poor prognosis of head and neck cancer
err2014-06-27
err0
PREAI
errJisoo Park; Jin‐man Kim; Jeong Kyu Park; Songmei Huang; Seo Young Kwak; Kyeung A. Ryu; Gyeyeong Kong; Jongsun Park; Bon Seok Koo
err分享
err收藏
err分享
err收藏
err分享
err收藏
Barrier to rotation about the N–N bond in 1,1′-bipiperidine
err1981-01-01
err0
PREAI
errKeiichiro Ogawa; Yoshito Takeuchi; Hiroshi Suzuki; Yujiro Nomura
err分享
err收藏
err分享
err收藏
学者 查看更多内容