返回
Machine learning in automated text categorization
DOI:10.1145/505282.505283.png)
摘要
En 中文
The automated categorization (or classification) of texts into predefined categories has witnessed a booming interest in the last 10 years, due to the increased availability of documents in digital form and the ensuing need to organize them. In the research community the dominant approach to this problem is based on machine learning techniques: a general inductive process automatically builds a classifier by learning, from a set of preclassified documents, the characteristics of the categories. The advantages of this approach over the knowledge engineering approach (consisting in the manual definition of a classifier by domain experts) are a very good effectiveness, considerable savings in terms of expert labor power, and straightforward portability to different domains. This survey discusses the main approaches to text categorization that fall within the machine learning paradigm. We will discuss in detail issues pertaining to three different problems, namely, document representation, classifier construction, and classifier evaluation.
Keyword:
algorithms
experimentation
theory
machine learning
text categorization
text classification
期刊
IF:
28
论文数:
2.4K
被引数:
3.5W
机构
暂无机构信息
引用论文
Enhancing the production of biogas through anaerobic co-digestion of agricultural waste and chemical pre-treatments
Chemosphere
IF0
Is this document relevant? ... probably: A survey of probabilistic models in information retrieval这份文件是否相关?...可能: 信息检索中的概率模型综述

