arrow
Return

EDGES: An Efficient Distributed Graph Embedding System on GPU Clusters

delete2020-01-01
delete5
PRE
AI
D
Dongxu Yang
J
Junhong Liu
J
Junjie Lai *
DOI:10.1109/TPDS.2020.3041219delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Graph embedding training models access parameters sparsely in a one-hot manner. Currently, the distributed graph embedding neural network is learned by data parallel with the parameter server, which suffers significant performance and scalability problems. In this article, we analyze the problems and characteristics of training this kind of models on distributed GPU clusters for the first time, and find that fixed model parameters scattered among different machine nodes are a major limiting factor for efficiency. Based on our observation, we develop an efficient distributed graph embedding system called EDGES, which can utilize GPU clusters to train large graph models with billions of nodes and trillions of edges using data and model parallelism. Within the system, we propose a novel dynamic partition architecture for training these models, achieving at least one half of communication reduction compared to existing training systems. According to our evaluations on real-world networks, our system delivers a competitive accuracy for the trained embeddings, and significantly accelerates the training process of the graph node embedding neural network, achieving a speedup of 7.23x and 18.6x over the existing fastest training system on single node and multi-node, respectively. As for the scalability, our experiments show that EDGES obtains a nearly linear speedup.
Keywords:
Training
Data models
Servers
Graphics processing units
Computational modeling
Neural networks
Training data
Large-scale distributed training
graph node embedding
GPU clusters
parallel algorithm
scalability
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Parallel and Distributed Systems cover
IEEE Transactions on Parallel and Distributed Systems
IF:
6
Papers:
5.2K
Citations:
1.1W

Organization

No organization information available
Cited Papers

Cited Papers

Prostate-derived Ets factor, an oncogenic driver in breast cancer
err2017-05-04
err0
errOAAI
errAshwani K Sood; Joseph Geradts; Jessica Young
errShare
errSave
Insurance activity and economic performance: Fresh evidence from asymmetric panel causality tests
err2018-10-24
err0
errOAAI
errAbdulnasser Hatemi‐J; Chi‐Chuan Lee; Chien‐Chiang Lee; Rangan Gupta
errShare
errSave
errShare
errSave
Induced Phase Transition in BiFeO3by High-Field Electron Spin Resonance
err2004-01-01
err0
PREAI
errD. VIEHLAND; J. F. LI; S. ZVYAGIN; A. P. PYATAKOV; A. BUSH; B. RUETTE; V. I. BELOTELOV; A. K. ZVEZDIN
errShare
errSave
researcher View more