返回
Data Lake Organization
DOI:10.1109/TKDE.2021.3091101.png)
摘要
En 中文
We consider the problem of building an organizational directory of data lakes to support effective user navigation. The organization directory is defined as an acyclic graph that contains nodes representing sets of attributes and edges indicating subset relationships between nodes. A probabilistic model is constructed to model user navigational behaviour. The model also predicts the likelihood of users finding relevant tables in a data lake given an organization. We formulate the data lake organization problem as an optimization over the organizational structure in order to maximize the expected likelihood of discovering tables by navigating. An approximation algorithm is proposed with an analysis of its error bound. The effectiveness and efficiency of the algorithm are evaluated on both synthetic and real data lakes. Our experiments show that our algorithm constructs organizations that outperform many existing organizations including an existing hand-curated taxonomy, a linkage graph, and a common baseline organization. We have also conducted a formal user study which shows that navigation can help users discover relevant tables that are not easily accessible by keyword search queries. This suggests that keyword search and navigation using an organization are complementary modalities for data discovery in data lakes.
Keyword:
Data lake
dataset discovery
taxonomy
structure learning
期刊
IF:
10.4
论文数:
6.8K
被引数:
3.2W
机构
引用论文
Being Bayesian about network structure. A Bayesian approach to structure discovery in Bayesian networks关于网络结构的贝叶斯。贝叶斯网络中结构发现的贝叶斯方法
MACHINE LEARNING
IF2.9

