arrow
返回

Roman urdu hate speech detection using hybrid machine learning models and hyperparameter optimization

delete2024-11-19
delete1
delete
OA
AI
W
Waqar Ashiq
S
Samra Kanwal
A
Adnan Rafique
M
Muhammad Waqas
T
Tahir Khurshaid *
E
Elizabeth Caro Montero
A
Alonso, Alicia Bustamante
I
Imran Ashraf *
DOI:10.1038/s41598-024-79106-7delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
With the rapid increase of users over social media, cyberbullying, and hate speech problems have arisen over the past years. Automatic hate speech detection (HSD) from text is an emerging research problem in natural language processing (NLP). Researchers developed various approaches to solve the automatic hate speech detection problem using different corpora in various languages, however, research on the Urdu language is rather scarce. This study aims to address the HSD task on Twitter using Roman Urdu text. The contribution of this research is the development of a hybrid model for Roman Urdu HSD, which has not been previously explored. The novel hybrid model integrates deep learning (DL) and transformer models for automatic feature extraction, combined with machine learning algorithms (MLAs) for classification. To further enhance model performance, we employ several hyperparameter optimization (HPO) techniques, including Grid Search (GS), Randomized Search (RS), and Bayesian Optimization with Gaussian Processes (BOGP). Evaluation is carried out on two publicly available benchmarks Roman Urdu corpora comprising HS-RU-20 corpus and RUHSOLD hate speech corpus. Results demonstrate that the Multilingual BERT (MBERT) feature learner, paired with a Support Vector Machine (SVM) classifier and optimized using RS, achieves state-of-the-art performance. On the HS-RU-20 corpus, this model attained an accuracy of 0.93 and an F1 score of 0.95 for the Neutral-Hostile classification task, and an accuracy of 0.89 with an F1 score of 0.88 for the Hate Speech-Offensive task. On the RUHSOLD corpus, the same model achieved an accuracy of 0.95 and an F1 score of 0.94 for the Coarse-grained task, alongside an accuracy of 0.87 and an F1 score of 0.84 for the Fine-grained task. These results demonstrate the effectiveness of our hybrid approach for Roman Urdu hate speech detection.
Keyword:
Hate speech detection
Deep learning
Model optimization
Urdu text classification
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Scientific Reports 封面图
Scientific Reports
IF:
3.9
论文数:
27.8W
被引数:
83.5W

机构

U
university of management & technology (umt)
学者数:
1.9K
论文数: 1.8K
被引数: 6
U
University of Tasmania
学者数:
1.3W
论文数: 1.3W
被引数: 1.8W
Y
Yeungnam University
学者数:
1.0W
论文数: 1.3W
被引数: 1.4W
学者 查看更多机构
引用论文

引用论文

err分享
err收藏
Hate Speech and Target Community Detection in Nastaliq Urdu Using Transfer Learning Techniques
err2024-01-01
err6
errOAAI
errMalik, Muhammad Shahid Iqbal; Nawaz, Aftab; Jamjoom, Mona Mamdouh
err分享
err收藏
Multi-label emotion classification of Urdu tweets
err2022-04-22
err20
errOAAI
errAshraf, Noman; Khan, Lal; Butt, Sabur; Chang, Hsien-Tsung; Sidorov, Grigori; Gelbukh, Alexander
err分享
err收藏
Separating the output ports of a Bragg interferometer via velocity selective transport
err2022-07-05
err0
errOAAI
errR. Piccon; S. Sarkar; J. Gomes Baptista; S. Merlet; F. Pereira Dos Santos
err分享
err收藏
学者 查看更多内容