arrow
返回

Optimizing distributed data stream processing by tracing

delete2019-01-01
delete9
PRE
AI
Z
Zoltán Zvara *
P
Péter G. N. Szabó
B
Barnabás Balázs
A
András A. Benczúr
DOI:10.1016/j.future.2018.06.047delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Heterogeneous mobile, sensor, loT, smart environment, and social networking applications have recently started to produce unbounded, fast, and massive-scale streams of data that have to be processed on the fly. Systems that process such data have to be enhanced with detection for operational exceptions and with triggers for both automated and manual operator actions. In this paper, we illustrate how tracing in distributed data processing systems can be applied to detecting changes in data and operational environment to maintain the efficiency of heterogeneous data stream processing systems under potentially changing data quality and distribution. By the tracing of individual input records, we can (1) identify outliers in a web crawling and document processing system and use the insights to define URL filtering rules; (2) identify heavy keys, such as NULL, that should be filtered before processing; (3) give hints to improve the key-based partitioning mechanisms; and (4) measure the limits of overpartitioning if heavy thread-unsafe libraries are imported. By using Apache Spark as illustration, we show how various data stream processing efficiency issues can be mitigated or optimized by our distributed tracing engine. We describe and qualitatively compare two different designs, one based on reporting to a distributed database and another based on trace piggybacking. Our prototype implementation consists of wrappers suitable for JVM environments in general, with minimal impact on the source code of the core system. Our tracing framework is the first to solve tracing in multiple systems across boundaries and to provide detailed performance measurements suitable for automated optimization, not just debugging. (C) 2018 Elsevier B.V. All rights reserved.
Keyword:
Distributed data processing
Data stream processing
Distributed tracing
Data provenance
Apache Spark
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
论文数:
6.8K
被引数:
2.3W

机构

H
Hungarian Academy of Sciences
学者数:
7.9K
论文数: 5.6K
被引数: 7.5K
引用论文

引用论文

err分享
err收藏
On evaluating stream learning algorithms
err2012-10-24
err360
errOAAI
errGama, Joao; Sebastiao, Raquel; Rodrigues, Pedro Pereira
err分享
err收藏
Power-Law Distributions in Empirical Data经验数据中的幂律分布
err2009-11-04
err6.7K
errOAAI
errClauset, Aaron; Shalizi, Cosma Rohilla; Newman, M. E. J.
err分享
err收藏
err分享
err收藏
Conserving and promoting evenness: organic farming and fire‐based wildland management as case studies
err2012-09-01
err0
errOAAI
errDavid W. Crowder; Tobin D. Northfield; Richard Gomulkiewicz; William E. Snyder
err分享
err收藏
没有更多内容