返回
A similarity-based approach for data stream classification
DOI:10.1016/j.eswa.2013.12.041.png)
摘要
En 中文
Incremental learning techniques have been used extensively to address the data stream classification problem. The most important issue is to maintain a balance between accuracy and efficiency, i.e., the algorithm should provide good classification performance with a reasonable time response. This work introduces a new technique, named Similarity-based Data Stream Classifier (SimC), which achieves good performance by introducing a novel insertion/removal policy that adapts quickly to the data tendency and maintains a representative, small set of examples and estimators that guarantees good classification rates. The methodology is also able to detect novel classes/labels, during the running phase, and to remove useless ones that do not add any value to the classification process. Statistical tests were used to evaluate the model performance, from two points of view: efficacy (classification rate) and efficiency (online response time). Five well-known techniques and sixteen data streams were compared, using the Friedman's test. Also, to find out which schemes were significantly different, the Nemenyi's, Holm's and Shaffer's tests were considered. The results show that SimC is very competitive in terms of (absolute and streaming) accuracy, and classification/updating time, in comparison to several of the most popular methods in the literature. (C) 2014 Elsevier Ltd. All rights reserved.
Keyword:
Data streams
Classification
Similarity
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
7.5
论文数:
3.0W
被引数:
10.2W
机构
引用论文
Tolerating concept and sampling shift in lazy learning using prediction error context switching使用预测误差上下文切换在懒惰学习中容忍概念和采样移位
Electrical Conductivity of Reproductive Tissue for Detection of Estrus in Dairy Cows用于检测奶牛发情的繁殖组织电导率

