arrow
返回

Pattern on demand in transactional distributed databases

delete2022-02-01
delete1
delete
OA
AI
L
Lamine Diop *
C
Cheikh Talibouya Diop
A
Arnaud Giacometti
A
Arnaud Soulet
DOI:10.1016/j.is.2021.101908delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Many applications rely on distributed databases like sensor networks or the Semantic Web. However, only few methods exist to extract patterns without centralizing the data by following the exhaustive extraction paradigm. Their principle is to extract a unique large collection of frequent patterns that will be used for all downstream applications. Unfortunately, the communication of this large collection from the different nodes is often more expensive than the database centralization. Furthermore, this rigid principle is not suited to modern data analysis where data and analyst needs change daily. It is both too expensive to repeat the exhaustive extraction for each change and it is not possible to build the ideal collection of patterns to meet all the needs. To circumvent this difficulty, this paper revisits the problem of pattern mining in distributed databases by adopting the Pattern-On-Demand paradigm. This principle consists in instantly extracting the patterns at the moment when the analyst needs them. Specifically, we propose a new pattern sampling algorithm, named DDSAMPLING, that randomly draws a pattern from a transactional distributed database with a probability proportional to its interest. We demonstrate the soundness of DDSAMPLING and analyze its time complexity. Finally, experiments on benchmark datasets highlight its low communication cost and its robustness against network and node failures. We also illustrate its interest on real-world data from the Semantic Web for detecting outlier entities in DBpedia and Wikidata. In addition, our output space sampling method is more parsimonious in terms of communication cost than a baseline relying on input space sampling. (C) 2021 Elsevier Ltd. All rights reserved.
Keyword:
Pattern mining
Knowledge base
Pattern on Demand
Pattern sampling
Outlier detection
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Enterprise Information Systems 封面图
Enterprise Information Systems
IF:
3.9
论文数:
2.8K
被引数:
1.8K

机构

U
universite de tours
学者数:
5.3K
论文数: 3.5K
被引数: 2
U
universite gaston berger
学者数:
294
论文数: 183
被引数: 4
引用论文

引用论文

err分享
err收藏
Computing maximal and minimal trap spaces of Boolean networks
err2015-10-07
err0
PREAI
errHannes Klarner; Alexander Bockmayr; Heike Siebert
err分享
err收藏
Genotypic Variability in Vulnerability of Leaf Xylem to Cavitation in Water-Stressed and Well-Irrigated Sugarcane
err1992-10-01
err0
errOAAI
errHoward S. Neufeld; David A. Grantz; Frederick C. Meinzer; Guillermo Goldstein; Gayle M. Crisosto; Carlos Crisosto
err分享
err收藏
err分享
err收藏
Wikidata: A Free Collaborative Knowledgebase
err2014-09-23
err1.9K
errOAAI
errVrandecic, Denny; Kroetzsch, Markus
err分享
err收藏
err分享
err收藏
Development of a pyrF-based counterselectable system for targeted gene deletion in Streptomyces rimosus
err2021-05-13
err0
errOAAI
errYiying Yang; Qingqing Sun; Yang Liu; Hanzhi Yin; Wenping Yang; Yang Wang; Ying Liu; Yuxian Li; Shen Pang; Wenxi Liu; Qian Zhang; Fang Yuan; Shiwen Qiu; Jiong Li; Xuefeng Wang; Keqiang Fan; Weishan Wang; Zilong Li; Shouliang Yin
err分享
err收藏
学者 查看更多内容