arrow
返回

Query-Driven Learning for Predictive Analytics of Data Subspace Cardinality

delete2017-06-29
delete23
PRE
AI
C
Christos Anagnostopoulos *
P
Peter Triantafillou
DOI:10.1145/3059177delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Fundamental to many predictive analytics tasks is the ability to estimate the cardinality (number of data items) of multi-dimensional data subspaces, defined by query selections over datasets. This is crucial for data analysts dealing with, e.g., interactive data subspace explorations, data subspace visualizations, and in query processing optimization. However, in many modern data systems, predictive analytics may be (i) too costly money-wise, e.g., in clouds, (ii) unreliable, e.g., in modern Big Data query engines, where accurate statistics are difficult to obtain/maintain, or (iii) infeasible, e.g., for privacy issues. We contribute a novel, query-driven, function estimation model of analyst-defined data subspace cardinality. The proposed estimation model is highly accurate in terms of prediction and accommodating the well-known selection queries: multi-dimensional range and distance-nearest neighbors (radius) queries. Our function estimation model: (i) quantizes the vectorial query space, by learning the analysts' access patterns over a data space, (ii) associates query vectors with their corresponding cardinalities of the analyst-defined data subspaces, (iii) abstracts and employs query vectorial similarity to predict the cardinality of an unseen/unexplored data subspace, and (iv) identifies and adapts to possible changes of the query subspaces based on the theory of optimal stopping. The proposed model is decentralized, facilitating the scaling-out of such predictive analytics queries. The research significance of the model lies in that (i) it is an attractive solution when data-driven statistical techniques are undesirable or infeasible, (ii) it offers a scale-out, decentralized training solution, (iii) it is applicable to different selection query types, and (iv) it offers a performance that is superior to that of data-driven approaches.
Keyword:
Predicive analytics
predictive learning
data subspace exploration
analytics selection queries
vector regression quantization
optimal stopping theory
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

ACM Transactions on Knowledge Discovery from Data 封面图
ACM Transactions on Knowledge Discovery from Data
IF:
4.8
论文数:
1.3K
被引数:
4.4K

机构

U
university of glasgow
学者数:
3.5W
论文数: 3.1W
被引数: 37
引用论文

引用论文

err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
err分享
err收藏
A rough path perspective on renormalization
err2019-12-01
err0
errOAAI
errY. Bruned; I. Chevyrev; P.K. Friz; R. Preiß
err分享
err收藏
err分享
err收藏
err分享
err收藏
Essentials of the self-organizing map
err2013-01-01
err1.1K
PREAI
errKohonen, Teuvo
err分享
err收藏
学者 查看更多内容