arrow
返回

Incremental dataflow execution, resource efficiency and probabilistic guarantees with Fuzzy Boolean nets

delete2015-05-01
delete3
PRE
AI
S
S. N. Esteves *
J
João Nuno Silva
J
João Paulo Carvalho
L
Luís Veiga
DOI:10.1016/j.jpdc.2015.03.001delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Currently, there is a strong need for organizations to analyze and process ever-increasing volumes of data in order to answer to real-time processing demands. Such continuous and data-intensive processing is often achieved through the composition of complex data-intensive workflows (i.e., dataflows). Dataflow management systems typically enforce strict temporal synchronization across the various processing steps. Non-synchronous behavior often has to be explicitly programmed on an ad-hoc basis, which requires additional lines of code in programs and thus the possibility of errors. More so, in a large set of scenarios for continuous and incremental processing, the output of dataflow applications at each execution can suffer almost no difference when comparing to the previous execution, and therefore resources, energy and computational power are unknowingly wasted. To face such lack of efficiency, transparency, and generality, we introduce the notion of Quality-of-Data (QoD), which describes the level of changes required on a data store that cause the triggering of processing steps. This, so that the dataflow (re-)execution is reduced until its outcome would reach a significant and meaningful variation, which is inside a specified freshness limit. Based on the QoD notion, we propose a novel dataflow model, with framework (Fluxy), for orchestrating data-intensive processing steps, which communicate data via a NoSQL storage, and whose triggering semantics is driven by dynamic QoD constraints automatically defined for different datasets by means of Fuzzy Boolean Nets. These nets give probabilistic guarantees about the prediction of the cumulative error between consecutive dataflow executions. With Fluxy, we demonstrate how dataflows can be leveraged to respond to quality boundaries (that can be seen as SLAs) to deliver controlled and augmented performance, rationalization of resources, and task prioritization. (C) 2015 Elsevier Inc. All rights reserved.
Keyword:
Workflow
Dataflow
Incremental processing
Continuous processing
Data-intensive
NoSQL
Quality-of-service
Machine learning
Fuzzy logic
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Journal of Parallel and Distributed Computing 封面图
Journal of Parallel and Distributed Computing
IF:
4
论文数:
3.8K
被引数:
4.8K

机构

I
inesc-id
学者数:
636
论文数: 504
被引数: 0
引用论文

引用论文

Polymorphisms in MMP-2 and TIMP-2 in Turkish patients with prostate cancer
err2014-01-01
err0
errOAAI
errKürşat Oğuz YAYKAŞLI; Muhammet Ali KAYIKÇI; Nesibe YAMAK; Hatice SOĞUKTAŞ; Selma DÜZENLİ; Ali Osman ARSLAN; Ahmet METİN; Ertuğrul KAYA; Ömer Faruk HATİPOĞLU
err分享
err收藏
Simulating scalar field theories on quantum computers with limited resources
err2023-03-03
err0
errOAAI
errAndy C. Y. Li; Alexandru Macridin; Stephen Mrenna; Panagiotis Spentzouris
err分享
err收藏
学者 查看更多内容