返回
PrIter: A Distributed Framework for Prioritizing Iterative Computations
DOI:10.1109/TPDS.2012.272.png)
摘要
En 中文
Iterative computations are pervasive among data analysis applications, including web search, online social network analysis, recommendation systems, and so on. These applications typically involve data sets of massive scale. Fast convergence of the iterative computations on the massive data set is essential for these applications. In this paper, we explore the opportunity for accelerating iterative computations by prioritization. Instead of performing computations on all data points without discrimination, we prioritize the computations that help convergence the most, so that the convergence speed of iterative process is significantly improved. We develop a distributed computing framework, PrIter, which supports the prioritized execution of iterative computations. PrIter either stores intermediate data in memory for fast convergence or stores intermediate data in files for scaling to larger data sets. We evaluate PrIter on a local cluster of machines as well as on Amazon EC2 Cloud. The results show that PrIter achieves up to 50 x speedup over Hadoop for a series of iterative algorithms. In addition, PrIter is shown better performance for iterative computations than other state-of-the-art distributed frameworks such as Spark and Piccolo.
Keyword:
PrIter
prioritized iteration
iterative algorithms
MapReduce
distributed framework
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6
论文数:
5.2K
被引数:
1.1W
机构
引用论文
Incidence of Dementia in Relation to Genetic Variants at PITX2, ZFHX3, and ApoE ε4 in Atrial Fibrillation Patients房颤患者中与PITX2,ZFHX3和apoeε4遗传变异相关的痴呆发生率
Affinity labeling of bovine carboxypeptidase A γLeu by N-bromoacetyl-N-methyl-L-phenylalanine. I. Kinetics of inactivation
Biochemistry
IF0
Enzymology with a Spin-Labeled Phospholipase C: Soluble Substrate Binding by 31P NMR from 0.005 to 11.7 T具有自旋标记的磷脂酶C的酶学: 通过 31P NMR从0.005到11.7 T的可溶性底物结合
Biochemistry
IF0

