返回
HDM: A Composable Framework for Big Data Processing
DOI:10.1109/TBDATA.2017.2690906.png)
摘要
En 中文
Over the past years, frameworks such as MapReduce and Spark have been introduced to ease the task of developing big data programs and applications. However, the jobs in these frameworks are roughly defined and packaged as executable jars without any functionality being exposed or described. This means that deployed jobs are not natively composable and reusable for subsequent development. Besides, it also hampers the ability for applying optimizations on the data flow of job sequences and pipelines. In this paper, we present the Hierarchically Distributed Data Matrix (HDM) which is a functional, strongly-typed data representation for writing composable big data applications. Along with HDM, a runtime framework is provided to support the execution, integration and management of HDM applications on distributed infrastructures. Based on the functional data dependency graph of HDM, multiple optimizations are applied to improve the performance of executing HDM jobs. The experimental results show that our optimizations can achieve improvements between 10 to 40 percent of the Job-Completion-Time for different types of applications when compared with the current state of art, Apache Spark.
Keyword:
Big data processing
parallel programming
functional programming
distributed systems
system architecture
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
I
IF:
5.7
论文数:
860
被引数:
3.0K
机构
引用论文
Capacity optimization of hybrid renewable energy system considering part-load ratio and resource endowment考虑部分负荷率和资源禀赋的混合可再生能源系统容量优化

