返回
Towards a Novel Framework for Automatic Big Data Detection
DOI:10.1109/ACCESS.2020.3030562.png)
摘要
En 中文
Big data is a relative concept. It is the combination of data, application, and platform properties. Recently, big data specific technologies have emerged, including software frameworks, databases, hardware accelerators, storage technologies, etc. However, the automatic selection of these solutions for big data computations remains a non-trivial task. Presently, the big data tools are selected by analyzing the problem manually, or by using several performance prediction techniques. The manual identification is based on the data properties only, whereas the performance predictors only estimate basic execution metrics without linking them with big data (3Vs) thresholds. Hence, both ways of identification are mostly incorrect, which can lead to inefficient use of 3Vs optimizations, resulting into global inefficiency, reduced system performance, increasing power consumption, requiring greater effort on the part of the programming team, and misallocation of the hardware resources required for the task. In this regard, a novel framework has been proposed for automatic detection of 3Vs (Volume, Velocity, Variety) of big data, using machine learning. The detection is done through static code features, data, and platform properties, leading to relevant tool selection, and code generation, with minimal overheads, lesser programmer interventions, higher usability, and portability. Instead of handling each application with big data specialized solutions, or manually identifying the 3Vs, the framework can automatically detect and link the 3Vs to the relevant optimizations. Several standard applications have been tested using the proposed framework. In the case of volume, the average detection accuracy is up to 97.8% for seen and 95.9% for unseen applications. In the case of velocity, the average detection accuracy is up to 97.3% for seen and 92.6%; for unseen applications. There is no margin of error in variety detection, as it has straightforward computations without any predictions. Furthermore, an airline recommendation system case study strengthens the effectiveness of the proposed approach.
Keyword:
Big Data
Feature extraction
Tools
Hardware
Software
Measurement
Optimization
Big data (3Vs)
detection
LLVM
machine learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
GFlink: An In-Memory Computing Architecture on Heterogeneous CPU-GPU Clusters for Big DataGFlink: 面向大数据的cpu-gpu异构集群内存计算架构
A comparison of forecasting models for the resource usage of MapReduce applicationsMapReduce应用资源使用预测模型的比较
NEUROCOMPUTING
IF6.5
Camptothecin exhibits topoisomerase1-independent KMT1A suppression and myogenic differentiation in alveolar rhabdomyosarcoma cells
Oncotarget
IF0

