arrow
返回

PathGraph: A Path Centric Graph Processing System

delete2016-10-01
delete26
PRE
AI
袁
袁平鹏 (Pingpeng Yuan) *
Ling Liu 封面图
Ling Liu (Ling Liu)
金
金海 (Hai Jin)
DOI:10.1109/TPDS.2016.2518664delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Large scale iterative graph computation presents an interesting systems challenge due to two well known problems: (1) the lack of access locality and (2) the lack of storage efficiency. This paper presents PathGraph, a system for improving iterative graph computation on graphs with billions of edges. First, we improve the memory and disk access locality for iterative computation algorithms on large graphs by modeling a large graph using a collection of tree-based partitions. This enables us to use path-centric computation rather than vertex-centric or edge-centric computation. For each tree partition, we re-label vertices using DFS in order to preserve consistency between the order of vertex ids and vertex order in the paths. Second, a compact storage that is optimized for iterative graph parallel computation is developed in the PathGraph system. Concretely, we employ delta-compression and store tree-based partitions in a DFS order. By clustering highly correlated paths together as tree based partitions, we maximize sequential access and minimize random access on storage media. Third but not the least, our path-centric computation model is implemented using a scatter/gather programming model. We parallel the iterative computation at partition tree level and perform sequential local updates for vertices in each tree partition to improve the convergence speed. To provide well balanced workloads among parallel threads at tree partition level, we introduce the concept of multiple stealing points based task queue to allow work stealings from multiple points in the task queue. We evaluate the effectiveness of PathGraph by comparing with recent representative graph processing systems such as GraphChi and X-Stream etc. Our experimental results show that our approach outperforms the two systems on a number of graph algorithms for both in-memory and out-of-core graphs. While our approach achieves better data balance and load balance, it also shows better speedup than the two systems with the growth of threads.
Keyword:
Graphs and networks
concurrent programming
graph algorithms
data storage representations
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Parallel and Distributed Systems 封面图
IEEE Transactions on Parallel and Distributed Systems
IF:
6
论文数:
5.2K
被引数:
1.1W

机构

U
university system of georgia
学者数:
7.3W
论文数: 6.6W
被引数: 101
引用论文

引用论文

err分享
err收藏
Persistence in Paradox
err2018-09-20
err0
PREAI
errMiguel Pina e Cunha; Stewart Clegg
err分享
err收藏
Fusion proton diagnostic for the C-2 field reversed configuration
err2014-08-18
err0
PREAI
errR. M. Magee; R. Clary; S. Korepanov; A. Smirnov; E. Garate; K. Knapp; A. Tkachev
err分享
err收藏
Haldane Quantum Spin Chains
err2003-01-27
err0
PREAI
errJean‐Pierre Renard; Louis‐Pierre Regnault; Michel Verdaguer
err分享
err收藏
Performance-Guaranteed Attitude Tracking Control for RLV: A Finite-Time PPC Approach
err2024-08-01
err0
PREAI
errZongyi Guo; David Henry; Xiyu Gu; Stephane Ygorra; Jérôme Cieslak; Jianguo Guo
err分享
err收藏
学者 查看更多内容