arrow
返回

Kappa Updated Ensemble for drifting data stream mining

delete2019-10-02
delete133
delete
OA
AI
C
Cano, Alberto *
B
Bartosz Krawczyk
DOI:10.1007/s10994-019-05840-zdelete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Learning from data streams in the presence of concept drift is among the biggest challenges of contemporary machine learning. Algorithms designed for such scenarios must take into an account the potentially unbounded size of data, its constantly changing nature, and the requirement for real-time processing. Ensemble approaches for data stream mining have gained significant popularity, due to their high predictive capabilities and effective mechanisms for alleviating concept drift. In this paper, we propose a new ensemble method named Kappa Updated Ensemble (KUE). It is a combination of online and block-based ensemble approaches that uses Kappa statistic for dynamic weighting and selection of base classifiers. In order to achieve a higher diversity among base learners, each of them is trained using a different subset of features and updated with new instances with given probability following a Poisson distribution. Furthermore, we update the ensemble with new classifiers only when they contribute positively to the improvement of the quality of the ensemble. Finally, each base classifier in KUE is capable of abstaining itself for taking a part in voting, thus increasing the overall robustness of KUE. An extensive experimental study shows that KUE is capable of outperforming state-of-the-art ensembles on standard and imbalanced drifting data streams while having a low computational complexity. Moreover, we analyze the use of Kappa versus accuracy to drive the criterion to select and update the classifiers, the contribution of the abstaining mechanism, the contribution of the diversification of classifiers, and the contribution of the hybrid architecture to update the classifiers in an online manner.
Keyword:
Machine learning
Data streams
Concept drift
Classification
Ensemble learning
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Machine Learning 封面图
Machine Learning
IF:
2.9
论文数:
2.7K
被引数:
3.4W

机构

V
Virginia Commonwealth University
学者数:
2.2W
论文数: 1.8W
被引数: 1.9W
引用论文

引用论文

A Survey on Ensemble Learning for Data Stream Classification面向数据流分类的集成学习研究综述
err2017-03-27
err378
PREAI
errGomes, Heitor Murilo; Barddal, Jean Paul; Enembreck, Fabricio; Bifet, Albert
err分享
err收藏
A survey on feature drift adaptation: Definition, benchmark, challenges and future directions
err2017-05-01
err77
errOAAI
errBarddal, Jean Paul; Gomes, Heitor Murilo; Enembreck, Fabricio; Pfahringer, Bernhard
err分享
err收藏
Assessment and comparison of three different air quality indices in China
err2017-07-04
err0
errOAAI
errYouping Li; Ya Tang; Zhongyu Fan; Hong Zhou; Zhengzheng Yang
err分享
err收藏
err分享
err收藏
Tailored conditions for controlled and fast growth of surface-grafted PNIPAM brushes
err2016-08-01
err0
PREAI
errA. Pomorska; K. Wolski; A. Puciul-Malinowska; S. Zapotoczny
err分享
err收藏
学者 查看更多内容