arrow
Return

Decision Trees for Mining Data Streams Based on the Gaussian Approximation

delete2014-01-01
delete129
PRE
AI
L
Leszek Rutkowski *
M
Maciej Jaworski
L
Lena Pietruczuk
P
Piotr Duda
DOI:10.1109/TKDE.2013.34delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Since the Hoeffding tree algorithm was proposed in the literature, decision trees became one of the most popular tools for mining data streams. The key point of constructing the decision tree is to determine the best attribute to split the considered node. Several methods to solve this problem were presented so far. However, they are either wrongly mathematically justified (e. g., in the Hoeffding tree algorithm) or time-consuming (e.g., in the McDiarmid tree algorithm). In this paper, we propose a new method which significantly outperforms the McDiarmid tree algorithm and has a solid mathematical basis. Our method ensures, with a high probability set by the user, that the best attribute chosen in the considered node using a finite data sample is the same as it would be in the case of the whole data stream.
Keywords:
Data steam
decision trees
information gain
Gaussian approximation

Journal

IEEE Transactions on Knowledge and Data Engineering cover
IEEE Transactions on Knowledge and Data Engineering
IF:
10.4
Papers:
6.7K
Citations:
3.2W

Organization

T
technical university czestochowa
Scholars:
1.3K
Papers: 1.5K
Citations: 1