arrow
Return

Elastic Scaling for Data Stream Processing

delete2014-06-01
delete168
PRE
AI
B
Buğra Gedik *
S
Scott Schneider
M
Martin Hirzel
K
Kun‐Lung Wu
DOI:10.1109/TPDS.2013.295delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This article addresses the profitability problem associated with auto-parallelization of general-purpose distributed data stream processing applications. Auto-parallelization involves locating regions in the application's data flow graph that can be replicated at run-time to apply data partitioning, in order to achieve scale. In order to make auto-parallelization effective in practice, the profitability question needs to be answered: How many parallel channels provide the best throughput? The answer to this question changes depending on the workload dynamics and resource availability at run-time. In this article, we propose an elastic auto-parallelization solution that can dynamically adjust the number of channels used to achieve high throughput without unnecessarily wasting resources. Most importantly, our solution can handle partitioned stateful operators via run-time state migration, which is fully transparent to the application developers. We provide an implementation and evaluation of the system on an industrial-strength data stream processing platform to validate our solution.
Keywords:
Data stream processing
parallelization
elasticity
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Transactions on Parallel and Distributed Systems cover
IEEE Transactions on Parallel and Distributed Systems
IF:
6
Papers:
5.2K
Citations:
1.1W

Organization

I
ihsan dogramaci bilkent university
Scholars:
3.6K
Papers: 3.5K
Citations: 8
I
international business machines (ibm)
Scholars:
5.7K
Papers: 4.5K
Citations: 4