返回
A survey on data preprocessing for data stream mining: Current status and future directions
DOI:10.1016/j.neucom.2017.01.078.png)
摘要
En 中文
Data preprocessing and reduction have become essential techniques in current knowledge discovery scenarios, dominated by increasingly large datasets. These methods aim at reducing the complexity inherent to real-world datasets, so that they can be easily processed by current data mining solutions. Advantages of such approaches include, among others, a faster and more precise learning process, and more understandable structure of raw data. However, in the context of data preprocessing techniques for data streams have a long road ahead of them, despite online learning is growing in importance thanks to the development of Internet and technologies for massive data collection. Throughout this survey, we summarize, categorize and analyze those contributions on data preprocessing that cope with streaming data. This work also takes into account the existing relationships between the different families of methods (feature and instance selection, and discretization). To enrich our study, we conduct thorough experiments using the most relevant contributions and present an analysis of their predictive performance, reduction rates, computational time, and memory usage. Finally, we offer general advices about existing data stream preprocessing algorithms, as well as discuss emerging future challenges to be faced in the domain of data stream preprocessing. (C) 2017 Elsevier B.V. All rights reserved.
Keyword:
Data mining
Data stream
Concept drift
Data preprocessing
Data reduction
Feature selection
Instance selection
Data discretization
Online learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6.5
论文数:
2.5W
被引数:
6.5W
机构
引用论文
Tolerating concept and sampling shift in lazy learning using prediction error context switching使用预测误差上下文切换在懒惰学习中容忍概念和采样移位
A study of statistical techniques and performance measures for genetics-based machine learning: accuracy and interpretability基于遗传学的机器学习的统计技术和性能度量的研究: 准确性和可解释性
SOFT COMPUTING
IF2.5

