arrow
Return

Cache-Based Multi-Query Optimization for Data-Intensive Scalable Computing Frameworks

delete2020-03-04
delete6
PRE
AI
P
Pietro Michiardi
D
Damiano Carra *
DOI:10.1007/s10796-020-09995-2delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In modern large-scale distributed systems, analytics jobs submitted by various users often share similar work, for example scanning and processing the same subset of data. Instead of optimizing jobs independently, which may result in redundant and wasteful processing, multi-query optimization techniques can be employed to save a considerable amount of cluster resources. In this work, we introduce a novel method combining in-memory cache primitives and multi-query optimization, to improve the efficiency of data-intensive, scalable computing frameworks. By careful selection and exploitation of common (sub)expressions, while satisfying memory constraints, our method transforms a batch of queries into a new, more efficient one which avoids unnecessary recomputations. To find feasible and efficient execution plans, our method uses a cost-based optimization formulation akin to the multiple-choice knapsack problem. Extensive experiments on a prototype implementation of our system show significant benefits of worksharing for both TPC-DS workloads and detailed micro-benchmarks.
Keywords:
Large scale computing
Caching
Optimization
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Information Systems Frontiers cover
Information Systems Frontiers
IF:
8.3
Papers:
2.0K
Citations:
6.5K

Organization

I
imt - institut mines-telecom
Scholars:
7.4K
Papers: 6.4K
Citations: 5
EURECOM cover
EURECOM
Scholars:
250
Papers: 240
Citations: 62