arrow
返回

STRETCH: Virtual Shared-Nothing Parallelism for Scalable and Elastic Stream Processing

delete2022-12-01
delete2
delete
OA
AI
V
Vincenzo Gulisano *
Y
Yiannis Nikolakopoulos
A
Alessandro V. Papadopoulos
M
Marina Papatriantafilou
P
Philippas Tsigas
DOI:10.1109/TPDS.2022.3181979delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Stream processing applications extract value from raw data through Directed Acyclic Graphs of data analysis tasks. Shared-nothing (SN) parallelism is the de-facto standard to scale stream processing applications. Given an application, SN parallelism ins9tantiates several copies of each analysis task, making each instance responsible for a dedicated portion of the overall analysis, and relies on dedicated queues to exchange data among connected instances. On the one hand, SN parallelism can scale the execution of applications both up and out since threads can run task instances within and across processes/nodes. On the other hand, its lack of sharing can cause unnecessary overheads and hinder the scaling up when threads operate on data that could be jointly accessed in shared memory. This trade-off motivated us in studying a way for stream processing applications to leverage shared memory and boost the scale up (before the scale out) while adhering to the widely-adopted and SN-based APIs for stream processing applications. We introduce STRETCH, a framework that maximizes the scale up and offers instantaneous elastic reconfigurations (without state transfer) for stream processing applications. We propose the concept of Virtual Shared-Nothing (VSN) parallelism and elasticity and provide formal definitions and correctness proofs for the semantics of the analysis tasks supported by STRETCH, showing they extend the ones found in common Stream Processing Engines. We also provide a fully implemented prototype and show that STRETCH's performance exceeds that of state-of-the-art frameworks such as Apache Flink and offers, to the best of our knowledge, unprecedented ultra-fast reconfigurations, taking less than 40 ms even when provisioning tens of new task instances.
Keyword:
Parallel processing
Watermarking
Task analysis
Semantics
Metadata
Standards
Prototypes
Stream processing
shared-nothing parallelism
shared-memory
elasticity
scalability

期刊

IEEE Transactions on Parallel and Distributed Systems 封面图
IEEE Transactions on Parallel and Distributed Systems
IF:
6
论文数:
5.2K
被引数:
1.1W

机构

C
chalmers university of technology
学者数:
1.5W
论文数: 1.6W
被引数: 10
M
Malardalen University
学者数:
1.2K
论文数: 1.5K
被引数: 3
引用论文

引用论文

err分享
err收藏
Metallocarboxypeptidases
err1960-02-01
err0
errOAAI
errJoseph E. Coleman; Bert L. Vallee
err分享
err收藏
A Catalog of Stream Processing Optimizations
err2014-03-01
err203
errOAAI
errHirzel, Martin; Soule, Robert; Schneider, Scott; Gedik, Bugra; Grimm, Robert
err分享
err收藏
err分享
err收藏
DNA extraction from urea-preserved blood or blood clots for use in PCR
err1995-02-01
err0
PREAI
errAnnette Gelhaus; Britta Urban; Claude Pirmez
err分享
err收藏
err分享
err收藏
Temperature control for thermomechanical ring rolling
err2023-06-27
err0
errOAAI
errRémi Lafarge; Alexander Brosius
err分享
err收藏
ScaleJoin: A Deterministic, Disjoint-Parallel and Skew-Resilient Stream Join
err2021-06-01
err16
PREAI
errGulisano, Vincenzo; Nikolakopoulos, Yiannis; Papatriantafilou, Marina; Tsigas, Philippas
err分享
err收藏
学者 查看更多内容