返回
Bayesian network model for semi-structured document classification
DOI:10.1016/j.ipm.2004.04.009.png)
摘要
En 中文
Recently, a new community has started to emerge around the development of new information research methods for searching and analyzing semi-structured and XML like documents. The goal is to handle both content and structural information, and to deal with different types of information content (text, image, etc.). We consider here the task of structured document classification. We propose a generative model able to handle both structure and content which is based on Bayesian networks. We then show how to transform this generative model into a discriminant classifier using the method of Fisher kernel. The model is then extended for dealing with different types of content information (here text and images). The model was tested on three databases: the classical webKB corpus composed of HTML pages, the new INEX corpus which has become a reference in the field of ad-hoc retrieval for XML documents, and a multimedia corpus of Web pages. (C) 2004 Elsevier Ltd. All rights reserved.
Keyword:
statistical learning
Bayesian networks
categorization
structured documents
XML
machine learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
I
IF:
6.9
论文数:
5.2K
被引数:
1.4W
机构
暂无机构信息
引用论文
The hierarchical hidden Markov model: Analysis and applications分层隐马尔可夫模型: 分析与应用
MACHINE LEARNING
IF2.9
没有更多内容

