arrow
返回

Scalable aggregation predictive analytics

delete2017-12-12
delete19
delete
OA
AI
C
Christos Anagnostopoulos *
F
Fotis Savva
P
Peter Triantafillou
DOI:10.1007/s10489-017-1093-ydelete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
We introduce a predictive modeling solution that provides high quality predictive analytics over aggregation queries in Big Data environments. Our predictive methodology is generally applicable in environments in which large-scale data owners may or may not restrict access to their data and allow only aggregation operators like COUNT to be executed over their data. In this context, our methodology is based on historical queries and their answers to accurately predict ad-hoc queries' answers. We focus on the widely used set-cardinality, i.e., COUNT, aggregation query, as COUNT is a fundamental operator for both internal data system optimizations and for aggregation-oriented data exploration and predictive analytics. We contribute a novel, query-driven Machine Learning (ML) model whose goals are to: (i) learn the query-answer space from past issued queries, (ii) associate the query space with local linear regression & associative function estimators, (iii) define query similarity, and (iv) predict the cardinality of the answer set of unseen incoming queries, referred to the Set Cardinality Prediction (SCP) problem. Our ML model incorporates incremental ML algorithms for ensuring high quality prediction results. The significance of contribution lies in that it (i) is the only query-driven solution applicable over general Big Data environments, which include restricted-access data, (ii) offers incremental learning adjusted for arriving ad-hoc queries, which is well suited for query-driven data exploration, and (iii) offers a performance (in terms of scalability, SCP accuracy, processing time, and memory requirements) that is superior to data-centric approaches. We provide a comprehensive performance evaluation of our model evaluating its sensitivity, scalability and efficiency for quality predictive analytics. In addition, we report on the development and incorporation of our ML model in Spark showing its superior performance compared to the Spark's COUNT method.
Keyword:
Query-driven predictive analytics
Predictive modeling
Aggregation operators
Set cardinality prediction
Regression vector quantization
Self-organizing maps
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Applied Intelligence 封面图
Applied Intelligence
IF:
3.5
论文数:
7.6K
被引数:
1.7W

机构

U
university of glasgow
学者数:
3.5W
论文数: 3.1W
被引数: 37
引用论文

引用论文

Menstrual disorders and month of birth
err2009-07-09
err0
PREAI
errP.H. Jongbloet; W.M. Kersemaekers; G.A. Zielhuis; A.L.M. Verbeek
err分享
err收藏
A rough path perspective on renormalization
err2019-12-01
err0
errOAAI
errY. Bruned; I. Chevyrev; P.K. Friz; R. Preiß
err分享
err收藏
Controlled synthesis of α,β‐difunctional poly(hexafluoropropylene oxide)
err2003-03-12
err0
PREAI
errJi‐Feng Ding; Martin Proudmore; Richard H. Mobbs; A. Rashid Khan; Colin Price; Colin Booth
err分享
err收藏
The Effect of Age, Race, and Sex on Social Cognitive Performance in Individuals With Schizophrenia
err2017-05-01
err0
errOAAI
errAmy E. Pinkham; Skylar Kelsven; Chrystyna Kouros; Philip D. Harvey; David L. Penn
err分享
err收藏
学者 查看更多内容