arrow
Return

Optimizing Cloud Data Lake Queries With a Balanced Coverage Plan

delete2024-01-01
delete2
PRE
AI
G
Grisha Weintraub *
E
Ehud Gudes
S
Shlomi Dolev
J
Jeffrey D. Ullman
DOI:10.1109/TCC.2023.3339208delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Cloud data lakes emerge as an inexpensive solution for storing very large amounts of data. The main idea is the separation of compute and storage layers. Thus, cheap cloud storage is used for storing the data, while compute engines are used for running analytics on this data in on-demand mode. However, to perform any computation on the data in this architecture, the data should be moved from the storage layer to the compute layer over the network for each calculation. Obviously, that hurts calculation performance and requires huge network bandwidth. In this paper, we study different approaches to improve query performance in a data lake architecture. We define an optimization problem that can provably speed up data lake queries. We prove that the problem is NP-hard and suggest heuristic approaches. Then, we demonstrate through the experiments that our approach is feasible and efficient (up to x30 query execution time improvement based on the TPC-H benchmark).
Keywords:
Big Data applications
Cloud computing
Costs
Measurement
Engines
Computer architecture
Standards
Cloud storage
data lakes
query optimization

Journal

I
IEEE Transactions on Cloud Computing
IF:
5
Papers:
1.8K
Citations:
4.3K

Organization

S
Stanford University
Scholars:
9.6W
Papers: 8.2W
Citations: 17.0W
B
ben gurion university
Scholars:
1.3W
Papers: 1.0W
Citations: 5