arrow
返回

Handling numeric attributes when comparing Bayesian network classifiers: does the discretization method matter?

delete2011-04-06
delete34
PRE
AI
M
M. Julia Flores
J
José A. Gámez
A
Ana María Martínez *
J
José M. Puerta
DOI:10.1007/s10489-011-0286-zdelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Within the framework of Bayesian networks (BNs), most classifiers assume that the variables involved are of a discrete nature, but this assumption rarely holds in real problems. Despite the loss of information discretization entails, it is a direct easy-to-use mechanism that can offer some benefits: sometimes discretization improves the run time for certain algorithms; it provides a reduction in the value set and then a reduction in the noise which might be present in the data; in other cases, there are some Bayesian methods that can only deal with discrete variables. Hence, even though there are many ways to deal with continuous variables other than discretization, it is still commonly used. This paper presents a study of the impact of using different discretization strategies on a set of representative BN classifiers, with a significant sample consisting of 26 datasets. For this comparison, we have chosen Naive Bayes (NB) together with several other semi-Naive Bayes classifiers: Tree-Augmented Naive Bayes (TAN), k-Dependence Bayesian (KDB), Aggregating One-Dependence Estimators (AODE) and Hybrid AODE (HAODE). Also, we have included an augmented Bayesian network created by using a hill climbing algorithm (BNHC). With this comparison we analyse to what extent the type of discretization method affects classifier performance in terms of accuracy and bias-variance discretization. Our main conclusion is that even if a discretization method produces different results for a particular dataset, it does not really have an effect when classifiers are being compared. That is, given a set of datasets, accuracy values might vary but the classifier ranking is generally maintained. This is a very useful outcome, assuming that the type of discretization applied is not decisive future experiments can be d times faster, d being the number of discretization methods considered.
Keyword:
Discretization
Bayesian classifiers
AODE
Naive Bayes

期刊

Applied Intelligence 封面图
Applied Intelligence
IF:
3.5
论文数:
7.6K
被引数:
1.7W

机构

U
Universidad de Castilla-La Mancha
学者数:
9.9K
论文数: 9.1K
被引数: 7
引用论文

引用论文

Not so naive Bayes: Aggregating one-dependence estimators
err2005-01-01
err518
errOAAI
errWebb, GI; Boughton, JR; Wang, ZH
err分享
err收藏
Comparative Repeat Profiling of Two Closely Related Conifers (Larix decidua and Larix kaempferi) Reveals High Genome Similarity With Only Few Fast-Evolving Satellite DNAs
err2021-07-12
err0
errOAAI
errTony Heitkam; Luise Schulte; Beatrice Weber; Susan Liedtke; Sarah Breitenbach; Anja Kögler; Kristin Morgenstern; Marie Brückner; Ute Tröber; Heino Wolf; Doris Krabel; Thomas Schmidt
err分享
err收藏
err分享
err收藏
Bayesian network classifiers贝叶斯网络分类器
err1997-01-01
err3.8K
errOAAI
errFriedman, N; Geiger, D; Goldszmidt, M
err分享
err收藏
学者 查看更多内容