arrow
返回

The Renoir Dataflow Platform: Efficient Data Processing without Complexity

delete2024-11-01
delete0
delete
OA
AI
L
Luca De Martini *
A
Alessandro Margara
G
Gianpaolo Cugola
M
Marco Donadoni
E
Edoardo Morassutto
DOI:10.1016/j.future.2024.06.018delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Today, data analysis drives the decision-making process in virtually every human activity. This demands for software platforms that offer simple programming abstractions to express data analysis tasks and that can execute them in an efficient and scalable way. State-of-the-art solutions range from low-level programming primitives, which give control to the developer about communication and resource usage, but require significant effort to develop and optimize new algorithms, to high-level platforms that hide most of the complexities of parallel and distributed processing, but often at the cost of reduced efficiency. To reconcile these requirements, we developed Renoir, a novel distributed data processing platform written in Rust. Renoir provides a high-level dataflow programming model as mainstream data processing systems. It supports static and streaming data, it enables data transformations, grouping, aggregation, iterative computations, and time-based analytics, and it provides all these features incurring in a low overhead. In this paper, we present the programming model and the implementation details of Renoir. We evaluate it under heterogeneous workloads. We compare it with state-of-the-art solutions for data analysis and high-performance computing, as well as alternative research products, which offer different programming abstractions and implementation strategies. Renoir programs are compact and easy to write: developers need not care about low-level concerns such as resource usage, data serialization, concurrency control, and communication. At the same time, Renoir consistently presents comparable or better performance than competing solutions, by a large margin in several scenarios. We conclude that Renoir offers a good tradeoff between simplicity and performance, allowing developers to easily express complex data analysis tasks and achieve high performance and scalability.
Keyword:
Data analytics
Distributed processing
Batch processing
Stream processing
Dataflow model
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

F
Future Generation Computer Systems-The International Journal of eScience
IF:
6.1
论文数:
6.8K
被引数:
2.3W

机构

P
Polytechnic University of Milan
学者数:
2.0W
论文数: 1.8W
被引数: 24
引用论文

引用论文

Metallocarboxypeptidases
err1960-02-01
err0
errOAAI
errJoseph E. Coleman; Bert L. Vallee
err分享
err收藏
AFLP markers reveal high clonal diversity and extreme longevity in four key arctic‐alpine species
err2011-11-10
err0
PREAI
errLUCIENNE C. De WITTE; GEORG F. J. ARMBRUSTER; LUDOVIC GIELLY; PIERRE TABERLET; JÜRG STÖCKLIN
err分享
err收藏
Pharmacokinetics of methysergide and its metabolite methylergometrine in man
err1986-01-01
err0
PREAI
errU. Bredberg; G. S. Eyjolfsdottir; L. Paalzow; P. Tfelt-Hansen; V. Tfelt-Hansen
err分享
err收藏
err分享
err收藏
A Model and Survey of Distributed Data-Intensive Systems
err2023-08-26
err4
errOAAI
errMargara, Alessandro; Cugola, Gianpaolo; Felicioni, Nicolo; Cilloni, Stefano
err分享
err收藏
Incremental, Iterative Data Processing with Timely Dataflow
err2016-09-22
err29
errOAAI
errMurray, Derek G.; McSherry, Frank; Isard, Michael; Isaacs, Rebecca; Barham, Paul; Abadi, Martin
err分享
err收藏
Apache Spark: A Unified Engine for Big Data ProcessingApache Spark: 用于大数据处理的统一引擎
err2016-10-28
err1.7K
PREAI
errZaharia, Matei; Xin, Reynold S.; Wendell, Patrick; Das, Tathagata; Armbrust, Michael; Dave, Ankur; Meng, Xiangrui; Rosen, Josh; Venkataraman, Shivaram; Franklin, Michael J.; Ghodsi, Ali; Gonzalez, Joseph; Shenker, Scott; Stoica, Ion
err分享
err收藏
学者 查看更多内容