返回
Efficient feature selection techniques for sentiment analysis
DOI:10.1007/s11042-019-08409-z.png)
摘要
En 中文
Sentiment analysis is a domain of study that focuses on identifying and classifying the ideas expressed in the form of text into positive, negative and neutral polarities. Feature selection is a crucial process in machine learning. In this paper, we aim to study the performance of different feature selection techniques for sentiment analysis. Term Frequency Inverse Document Frequency (TF-IDF) is used as the feature extraction technique for creating feature vocabulary. Various Feature Selection (FS) techniques are experimented to select the best set of features from feature vocabulary. The selected features are trained using different machine learning classifiers Logistic Regression (LR), Support Vector Machines (SVM), Decision Tree (DT) and Naive Bayes (NB). Ensemble techniques Bagging and Random Subspace are applied on classifiers to enhance the performance on sentiment analysis. We show that, when the best FS techniques are trained using ensemble methods achieve remarkable results on sentiment analysis. We also compare the performance of FS methods trained using Bagging, Random Subspace with varied neural network architectures. We show that FS techniques trained using ensemble classifiers outperform neural networks requiring significantly less training time and parameters thereby eliminating the need for extensive hyper-parameter tuning.
Keyword:
Feature selection
Ensemble techniques
Sentiment analysis
Machine learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3
论文数:
2.0W
被引数:
3.2W
机构
引用论文
Clinical Efficacy of Recombinant Human Endostatin Combined With Cisplatin in the Treatment of Pleural Effusion of Lung Cancer重组人血管内皮抑制素联合顺铂治疗肺癌胸腔积液的临床疗效
Serial changes of prefrontal lobe growth in the patients with benign childhood epilepsy with centrotemporal spikes presenting with cognitive impairments/behavioral problems伴有认知障碍/行为问题的良性儿童癫痫患者中前额叶生长的系列变化


