arrow
返回

Probabilistic data structures for big data analytics: A comprehensive review

delete2020-01-01
delete33
PRE
AI
A
Amritpal Singh
S
Sahil Garg
R
Ravneet Kaur
S
Shalini Batra
N
Neeraj Kumar *
A
Albert Y. Zomaya
DOI:10.1016/j.knosys.2019.104987delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
An exponential increase in the data generation resources is widely observed in last decade, because of evolution in technologies such as-cloud computing, IoT, social networking, etc. This enormous and unlimited growth of data has led to a paradigm shift in storage and retrieval patterns from traditional data structures to Probabilistic Data Structures (PDS). PDS are a group of data structures that are extremely useful for Big data and streaming applications in order to avoid high-latency analytical processes. These data structures use hash functions to compactly represent a set of items in stream-based computing while providing approximations with error bounds so that well-formed approximations get built into data collections directly. Compared to traditional data structures, PDS use much less memory and constant time in processing complex queries. This paper provides a detailed discussion of various issues which are normally encountered in massive data sets such as-storage, retrieval, query,etc. Further, role of PDS in solving these issues is also discussed where these data structures are used as temporary accumulators in query processing. Several variants of existing PDS along with their application areas have also been explored which give a holistic view of domains where these data structures can be applied for efficient storage and retrieval of massive data sets. Mathematical proofs of various parameters considered in the PDS have also been discussed in the paper. Moreover, the relative comparison of various PDS with respect to various parameters is also explored. (C) 2019 Elsevier B.V. All rights reserved.
Keyword:
Big data
Internet of things (IoT)
Probabilistic data structures
Bloom filter
Quotient filter
Count min sketch
HyperLogLog counter
MM-hash
Locality sensitive hashing
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

K
Knowledge-Based Systems
IF:
7.6
论文数:
1.2W
被引数:
4.5W

机构

U
University of Sydney
学者数:
6.5W
论文数: 6.2W
被引数: 90
引用论文

引用论文

err分享
err收藏
err分享
err收藏
The Deletable Bloom Filter: A New Member of the Bloom Family
err2010-06-01
err54
errOAAI
errRothenberg, Christian Esteve; Macapuna, Carlos A. B.; Verdi, Fabio L.; Magalhaes, Mauricio F.
err分享
err收藏
err分享
err收藏
学者 查看更多内容