arrow
Return

A novel probabilistic feature selection method for text classification

delete2012-12-01
delete222
PRE
AI
A
Alper Kürşat Uysal *
S
Serkan Günal
DOI:10.1016/j.knosys.2012.06.005delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
High dimensionality of the feature space is one of the most important concerns in text classification problems due to processing time and accuracy considerations. Selection of distinctive features is therefore essential for text classification. This study proposes a novel filter based probabilistic feature selection method, namely distinguishing feature selector (DFS), for text classification. The proposed method is compared with well-known filter approaches including chi square, information gain, Gini index and deviation from Poisson distribution. The comparison is carried out for different datasets, classification algorithms, and success measures. Experimental results explicitly indicate that DFS offers a competitive performance with respect to the abovementioned approaches in terms of classification accuracy, dimension reduction rate and processing time. (C) 2012 Elsevier B.V. All rights reserved.
Keywords:
Feature selection
Filter
Pattern recognition
Text classification
Dimension reduction
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

A
Anadolu University
Scholars:
1.7K
Papers: 1.9K
Citations: 1.7K