arrow
Return

A Bayesian feature selection paradigm for text classification

delete2012-03-01
delete34
PRE
AI
G
Guozhong Feng
J
Jianhua Guo *
B
Bing‐Yi Jing
DOI:10.1016/j.ipm.2011.08.002delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The automated classification of texts into predefined categories has witnessed a booming interest, due to the increased availability of documents in digital form and the ensuing need to organize them. An important problem for text classification is feature selection, whose goals are to improve classification effectiveness, computational efficiency, or both. Due to categorization unbalancedness and feature sparsity in social text collection, filter methods may work poorly. In this paper, we perform feature selection in the training process, automatically selecting the best feature subset by learning, from a set of preclassified documents, the characteristics of the categories. We propose a generative probabilistic model, describing categories by distributions, handling the feature selection problem by introducing a binary exclusion/inclusion latent vector, which is updated via an efficient Metropolis search. Real-life examples illustrate the effectiveness of the approach. (C) 2011 Elsevier Ltd. All rights reserved.
Keywords:
Bayesian feature selection
Metropolis search
Mixture model
Text classification
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

I
Information Processing and Management
IF:
6.9
Papers:
5.2K
Citations:
1.4W

Organization

N
northeast normal university - china
Scholars:
1.2W
Papers: 9.2K
Citations: 23