返回
A logistic regression-based smoothing method for Chinese text categorization
DOI:10.1016/j.eswa.2011.03.036.png)
摘要
En 中文
Automatic Chinese text classification is an important and a well-known technology in the field of machine learning. The first step for solving Chinese text categorization problems is to tokenize the Chinese words from a sequence of non-segmented sentences. However, previous literatures often employ a Chinese word tokenizer that was trained with different sources and then perform the conventional text classification approaches. However, these taggers are not perfect and often provide incorrect word boundary information. In this paper, we propose an N-gram-based language model which takes word relations into account for Chinese text categorization without Chinese word tokenizer. To prevent from out-of-vocabulary, we also propose a novel smoothing approach based on logistic regression to improve accuracy. The experimental result shows that our approach outperforms traditional methods at least 11% on micro-average F-measure. (C) 2011 Elsevier Ltd. All rights reserved.
Keyword:
Text classification
N-gram-based classification
Feature selection
Word segmentation
Logistic regression
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
7.5
论文数:
3.0W
被引数:
10.2W
机构
引用论文
On machine learning methods for Chinese document categorization面向中文文档分类的机器学习方法研究
APPLIED INTELLIGENCE
IF3.5


