arrow
Return

Accelerating Large-Scale Heterogeneous Interaction Graph Embedding Learning via Importance Sampling

delete2020-12-07
delete10
delete
OA
AI
Y
Yugang Ji *
M
Mingyang Yin
H
Hongxia Yang
J
Jingren Zhou
V
Vincent W. Zheng
C
Chuan Shi
Y
Yuan Fang
DOI:10.1145/3418684delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
In real-world problems, heterogeneous entities are often related to each other through multiple interactions, forming a Heterogeneous Interaction Graph (HIG). While modeling HIGs to deal with fundamental tasks, graph neural networks present an attractive opportunity that can make full use of the heterogeneity and rich semantic information by aggregating and propagating information from different types of neighborhoods. However, learning on such complex graphs, often with millions or billions of nodes, edges, and various attributes, could suffer from expensive time cost and high memory consumption. In this article, we attempt to accelerate representation learning on large-scale HIGs by adopting the importance sampling of heterogeneous neighborhoods in a batch-wise manner, which naturally fits with most batch-based optimizations. Distinct from traditional homogeneous strategies neglecting semantic types of nodes and edges, to handle the rich heterogeneous semantics within HIGs, we devise both type-dependent and type-fusion samplers where the former respectively samples neighborhoods of each type and the latter jointly samples from candidates of all types. Furthermore, to overcome the imbalance between the down-sampled and the original information, we respectively propose heterogeneous estimators including the self-normalized and the adaptive estimators to improve the robustness of our sampling strategies. Finally, we evaluate the performance of our models for node classification and link prediction on five real-world datasets, respectively. The empirical results demonstrate that our approach performs significantly better than other state-of-the-art alternatives, and is able to reduce the number of edges in computation by up to 93%, the memory cost by up to 92% and the time cost by up to 86%.
Keywords:
Heterogeneous interaction graphs
large-scale graphs
type-dependent sampler
type-fusion sampler
importance sampling
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

ACM Transactions on Knowledge Discovery from Data cover
ACM Transactions on Knowledge Discovery from Data
IF:
4.8
Papers:
1.3K
Citations:
4.4K

Organization

A
alibaba group
Scholars:
1.1K
Papers: 789
Citations: 0
B
beijing university of posts & telecommunications
Scholars:
1.4W
Papers: 1.2W
Citations: 9
S
Singapore Management University
Scholars:
1.5K
Papers: 2.5K
Citations: 3.5K
researcher View more organizations