arrow
返回

On Two-Stage Feature Selection Methods for Text Classification

delete2018-01-01
delete35
delete
OA
AI
A
Alper Kürşat Uysal *
DOI:10.1109/ACCESS.2018.2863547delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Text classification is a high dimensional pattern recognition problem where feature selection is an important step. Although researchers still propose new feature selection methods, there exist many two-stage feature selection methods combining existing filter-based feature selection methods with feature transformation and wrapper-based feature selection methods in different ways. The main focus of the study is to extensively analyze two-stage feature selection methods for text classification from a different point of view. Two-stage feature selection methods that are constituted by combining filter-based local feature selection methods with feature transformation and wrapper-based feature selection methods were investigated in this paper. In the first stage, four different filter-based local feature selection methods and three different feature set construction methods were employed. Feature sets were constructed either by using maximum globalization policy (MAX), by using weighted averaging globalization policy (AVG), or by selecting an equal number of features for each class (EQ). In the second stage, principal component analysis (PCA), latent semantic indexing (LSI), or genetic algorithms were utilized. Various settings were evaluated with a linear support vector machines classifier on two benchmark data sets, namely, Reuters and Ohsumed using Micro-Fl and Macro-Fl scores. According to the findings, AVG and EQ feature set construction methods are usually more successful than MAX method for two-stage feature selection methods. Most of the highest accuracies were obtained by employing PCA feature transformation in the second stage. However, there is a strong linear correlation between PCA and LSI for all settings but the degree of correlation is slightly more for Ohsumed data set in comparison with the Reuters data set.
Keyword:
Feature selection
genetic algorithms
LSI
PCA
text classification
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

A
Anadolu University
学者数:
1.7K
论文数: 1.9K
被引数: 1.7K
引用论文

引用论文

Classification of sentiment reviews using n-gram machine learning approach
err2016-09-01
err330
PREAI
errTripathy, Abinash; Agrawal, Ankit; Rath, Santanu Kumar
err分享
err收藏
Relative discrimination criterion - A novel feature ranking method for text data
err2015-05-01
err56
PREAI
errRehman, Abdur; Javed, Kashif; Babri, Haroon A.; Saeed, Mehreen
err分享
err收藏
Text normalization and semantic indexing to enhance Instant Messaging and SMS spam filtering
err2016-09-01
err64
PREAI
errAlmeida, Tiago A.; Silva, Tiago P.; Santos, Igor; Gomez Hidalgo, Jose M.
err分享
err收藏
Authorship identification from unstructured texts
err2014-08-01
err51
PREAI
errZhang, Chunxia; Wu, Xindong; Niu, Zhendong; Ding, Wei
err分享
err收藏
学者 查看更多内容