arrow
Return

A discriminative and semantic feature selection method for text categorization

delete2015-07-01
delete49
PRE
AI
W
Wei Zong
F
Feng Wu
L
Lap-Keung Chu *
D
D. Sculli
DOI:10.1016/j.ijpe.2014.12.035delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Text categorization is an important and critical task in the current era of high volume data storage and handling. Feature selection is obviously one of the most important steps in text categorization. Traditional feature selection methods tend to only consider the correlation between features and categories, and have in the main ignored the semantic similarity between features and documents. To further explore this issue, this paper proposes a novel feature selection method that first selects features in documents with discriminative power and then computes the semantic similarity between features and documents. The proposed feature selection method is tested using a support vector machine (SVM) classifier upon two published datasets, viz. Reuters-21578 and 20-Newsgroups. The experimental results show that the proposed feature selection method generally outperforms the traditional feature selection methods for text categorization for both published datasets. (C) 2015 Elsevier B.V. All rights reserved.
Keywords:
Feature selection
Big data
Discriminative power
Semantic similarity
Text categorization
Support vector machine (SVM)
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

International Journal of Production Economics cover
International Journal of Production Economics
IF:
10
Papers:
7.9K
Citations:
3.6W

Organization

U
University of Hong Kong
Scholars:
4.1W
Papers: 3.9W
Citations: 10.1W
X
xi'an jiaotong university
Scholars:
9.2W
Papers: 6.6W
Citations: 75