arrow
返回

BERT-based ensemble learning for multi-aspect hate speech detection

delete2023-01-03
delete23
PRE
AI
A
Ahmed Cherif Mazari *
N
Nesrine Boudoukhani
A
Abdelhamid Djeffal
DOI:10.1007/s10586-022-03956-xdelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The social media world nowadays is overwhelmed with unfiltered content ranging from cyberbullying and cyberstalking to hate speech. Therefore, identifying and cleaning up such toxic language presents a big challenge and an active area of research. This study is dedicated to multi-aspect hate speech detection based on classifying text in multi-labels including 'identity hate', 'threat', 'insult', 'obscene', 'toxic' and 'severe toxic'. The proposed approach is based on the pre-trained Bidirectional Encoder Representations from Transformers (BERT) model combined with Deep Learning (DL) models to compose several ensemble learning architectures. The DL models used are built by stacking Bidirectional Long-Short Term Memory (Bi-LSTM) and/or Bidirectional Gated Recurrent Unit (Bi-GRU) on GloVe and FastText word embeddings. Whereby, these models and BERT are trained individually on multi-label hateful dataset and used in combination for hate speech detection tasks on social media. Thus, we demonstrate that encoding texts by using recent word embedding techniques as FastText and GloVe alongside Bi-LSTM and Bi-GRU can create models that, when combined with BERT, can enhance the ROC-AUC score to 98.63%.
Keyword:
BERT model
Multi-aspect hate speech
Language toxicity
Deep learning
Ensemble learning

期刊

C
Cluster Computing-The Journal of Networks Software Tools and Applications
IF:
4.1
论文数:
5.1K
被引数:
7.5K

机构

U
universite yahia fares medea
学者数:
410
论文数: 326
被引数: 5
U
universite mohamed khider biskra
学者数:
1.2K
论文数: 831
被引数: 1
引用论文

引用论文

err分享
err收藏
Misogyny Detection in Twitter: a Multilingual and Cross-Domain Study
err2020-11-01
err73
PREAI
errPamungkas, Endang Wahyu; Basile, Valerio; Patti, Viviana
err分享
err收藏
学者 查看更多内容