arrow
返回

Multi-Node Acceleration for Large-Scale GCNs

delete2022-01-01
delete5
delete
OA
AI
G
Gongjian Sun *
M
Mingyu Yan
D
Duo Wang
H
Han Li
W
Wenming Li
X
Xiaochun Ye
D
Dongrui Fan
Y
Yuan Xie
DOI:10.1109/TC.2022.3207127delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Limited by the memory capacity and computation power, singe-node graph convolutional neural network (GCN) accelerators cannot complete the execution of GCNs within a reasonable amount of time, due to the explosive size of graphs nowadays. Thus, large-scale GCNs call for a multi-node acceleration system (MultiAccSys) like tensor processing unit (TPU) Pod for large-scale neural network. In this work, we aim to scale up single-node GCN accelerator to accelerate GCNs on large-scale graphs. We first identify the communication pattern and challenges of multi-node acceleration for GCNs on large-scale graphs. We observe that (1) irregular coarse-grained communication patterns exist in the execution of GCNs in MultiAccSys, which introduces massive amount of redundant network transmissions and off-chip memory accesses; (2) the acceleration of GCNs in MultiAccSys is mainly bounded by network bandwidth but tolerates network latency. Guided by the above observations, we then propose MultiGCN, an efficient MultiAccSys for large-scale GCNs that trades network latency for network bandwidth. Specifically, by leveraging the network latency tolerance, we first propose a topology-aware multicast mechanism with a one put per multicast message-passing model to reduce transmissions and alleviate network bandwidth requirements. Second, we introduce a scatter-based round execution mechanism which cooperates with the multicast mechanism and reduces redundant off-chip memory accesses. Compared to the baseline MultiAccSys, MultiGCN achieves 4 & SIM; 12x speedup using only 28%$\sim$& SIM;68% energy, while reducing 32% transmissions and 73% off-chip memory accesses on average. Besides, MultiGCN not only achieves 2.5 & SIM; 8x speedup over the state-of-the-art multi-GPU solution, but also scales to large-scale graph as opposed to single-node GCN accelerators.
Keyword:
Deep learning
graph neural network
hardware accelerator
multi-node system
communication optimization

期刊

IEEE Transactions on Computers 封面图
IEEE Transactions on Computers
IF:
3.8
论文数:
5.4K
被引数:
9.8K

机构

C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
引用论文

引用论文

err分享
err收藏
err分享
err收藏
Landscape — a usefully ambiguous concept
err2014-11-14
err0
PREAI
errCHRIS GOSDEN; LESLEY HEAD
err分享
err收藏
Current observations in the southern Yellow Sea in summer
err2004-09-01
err0
PREAI
errTang Xiaohui; Wang Fan; Chen Yongli; Bai Hong; Hu Dunxin
err分享
err收藏
学者 查看更多内容