arrow
返回

Novel method for optimizing performance in resource constrained distributed data streams

delete2022-02-16
delete0
delete
OA
AI
R
Rashi Bhalla *
R
Russel Pears
M
M. Asif Naeem
F
Farhaan Mirza
DOI:10.1007/s10489-021-03019-5delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The Big Data Era has presented many opportunities for using data mining techniques to discover knowledge patterns across large and diverse collections of data where the volume of data is growing at an exponential rate. Recent approaches to Distributed Data Mining (DDM) have focused on addressing the heterogeneous nature of data sources. However, such approaches do not prioritize the reduction of data communication costs which could be prohibitive in large scale sensor networks where bandwidth is a limited resource. In fact, higher communication and computational costs are the two most prominent problems that have been encountered in heterogeneous distributed environments. Moreover, an effort to decrease the communications load in the distributed environment has an adverse influence on the classification accuracy. Therefore, the research challenge lies in maintaining a balance between transmission cost, computational cost, and accuracy. This paper proposes an algorithm Performance Optimizer in Distributed Stream Mining (PODSM) based on Bayesian Inference to reduce the communication volume and resource time in a heterogeneous distributed data mining environment while retaining prediction accuracy. The approach used in this work exploits the past data for calculating statistics and these statistics are then utilized for the new data. In other words, it imparts the ability to learn from experiences. As a result, our experimental evaluation reveals that a significant reduction in the communication load and an improvement in classification response time can be achieved across a diverse range of dataset types. Reduction of 34.66% was obtained with regard to communication overhead for one of the datasets with huge savings of nearly 27% in resource time. Importantly, instead of showing a negative effect on accuracy, this dataset observes an increment of 0.44% in accuracy.
Keyword:
Big data
Bayesian inference
Distributed data stream mining
Heterogeneous distributed data

期刊

Applied Intelligence 封面图
Applied Intelligence
IF:
3.5
论文数:
7.6K
被引数:
1.7W

机构

A
Auckland University of Technology
学者数:
4.0K
论文数: 4.4K
被引数: 4.7K
引用论文

引用论文

Ensemble learning for data stream analysis: A survey用于数据流分析的集成学习: 综述
err2017-09-01
err672
errOAAI
errKrawczyk, Bartosz; Minku, Leandro L.; Gama, Joao; Stefanowski, Jerzy; Wozniak, Michal
err分享
err收藏
err分享
err收藏
err分享
err收藏
Bayesian network classifiers贝叶斯网络分类器
err1997-01-01
err3.8K
errOAAI
errFriedman, N; Geiger, D; Goldszmidt, M
err分享
err收藏
Mechanisms of action of nitrates
err1994-10-01
err0
PREAI
errKristina E. Torfg�rd; Johan Ahlner
err分享
err收藏
Adaptive random forests for evolving data stream classification
err2017-06-13
err483
errOAAI
errGomes, Heitor M.; Bifet, Albert; Read, Jesse; Barddal, Jean Paul; Enembreck, Fabricio; Pfharinger, Bernhard; Holmes, Geoff; Abdessalem, Talel
err分享
err收藏
学者 查看更多内容