arrow
返回

A Stack-Centric Processing Model for Iterative Processing

delete2018-01-01
delete0
PRE
AI
Z
Zhifei Pang
伍
伍赛 (Sai Wu) *
陈
陈刚 (Gang Chen)
陈
陈珂 (Ke Chen)
Bingsheng He 封面图
Bingsheng He (Bingsheng He)
DOI:10.1109/TBDATA.2018.2841363delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Complex data mining algorithms are processed in multiple iterations, where output of one iteration is used as input for the subsequent iterations. Existing parallel programming frameworks, e.g., MapReduce, Pregel and Spark, adopt the breadth first search (BFS) strategy to process those iterative jobs. They invoke the user-defined functions for every key-value pair or vertex to produce all possible intermediate results for the next iteration. Such BFS strategy incurs high I/O overheads, because normally, the size of intermediate search results of BFS is exponential to the size of original data, making it impossible to maintain those intermediate results in memory. In this paper, we present a new type of parallel programming model, the stack-centric model, where all computations are defined for a stack maintained in the distributed shared memory. The stack can be adaptively split into multiple stacks and disseminated to different compute nodes for parallel processing. The most distinguished feature of the stack-centric model is its support for the depth first search (DFS) algorithm which incurs much less memory overhead than its BFS counterpart. The maximal memory usage of DFS algorithm is determined by the height of its search tree, and hence, it is possible to conduct the computation of DFS algorithm mostly in memory. Our stack-centric model is not a pure DFS framework. It supports the hybrid BFS and DFS algorithms by tuning the trade-off between memory usage and parallelism. To show the advantages of stack-centric model, we implement two algorithms, frequent pattern mining algorithm and DNA sequence matching algorithm, on both stack-centric model and Spark. The memory usage of stack-centric model is 10 times less than the Spark, resulting in a significant performance improvement.
Keyword:
Computational modeling
Data mining
Big Data
Sparks
Parallel processing
Vegetation
Indexes
Stack
iterative processing
data mining
DFS
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

I
IEEE Transactions on Big Data
IF:
5.7
论文数:
881
被引数:
3.0K

机构

N
National University of Singapore
学者数:
7.6W
论文数: 6.5W
被引数: 11.4W
Z
zhejiang university
学者数:
17.7W
论文数: 12.1W
被引数: 152
引用论文

引用论文

err分享
err收藏
In Vitro and in Silico Insights on the Biological Activities, Phenolic Compounds Composition of Hypericum perforatum L. Hairy Root Cultures
err2023-01-01
err0
errOAAI
errOliver Tusevski; Marija Todorovska; Jasmina Petreska Stanoeva; Marina Stefova; Sonja Gadzovska Simic
err分享
err收藏
Developing an Online Curriculum
err
IF0
err2004-01-01
err0
PREAI
errLynnette R. Porter
err分享
err收藏
Haldane Quantum Spin Chains
err2003-01-27
err0
PREAI
errJean‐Pierre Renard; Louis‐Pierre Regnault; Michel Verdaguer
err分享
err收藏
err分享
err收藏
学者 查看更多内容