返回
Exploiting structural information for semi-structured document categorization
DOI:10.1016/j.ipm.2005.06.003.png)
摘要
En 中文
This paper examines several different approaches to exploiting structural information in semi-structured document categorization. The methods under consideration are designed for categorization of documents consisting of a collection of fields, or arbitrary tree-structured documents that can be adequately modeled with such a flat structure. The approaches range from trivial modifications of text modeling to more elaborate schemes, specifically tailored to structured documents. We combine these methods with three different text classification algorithms and evaluate their performance on four standard datasets containing different types of semi-structured documents. The best results were obtained with stacking, an approach in which predictions based on different structural components are combined by a meta classifier. A further improvement of this method is achieved by including the flat text model in the final prediction. (c) 2005 Elsevier Ltd. All rights reserved.
Keyword:
text categorization
semi-structured documents
document structure
stacked generalization
support vector machines
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
I
IF:
6.9
论文数:
5.2K
被引数:
1.4W
机构
暂无机构信息
引用论文
没有更多内容

