arrow
返回

HHSD: Hindi Hate Speech Detection Leveraging Multi-Task Learning

delete2023-01-01
delete3
delete
OA
AI
P
Prashant Kapil *
G
Gitanjali Kumari
A
Asif Ekbal
S
Santanu Pal
A
Arindam Chatterjee
B
B. N. Vinutha
DOI:10.1109/ACCESS.2023.3312993delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Hate speech is now a frequent occurrence on social media. Recently, the majority of study was devoted to identifying hate speech in languages with abundant resources (e.g., English). However, relatively few works are developed for languages with limited resources (e.g., Hindi, the third most widely used language on earth). In this study, Hindi Hate Speech Dataset (HHSD) is created following a novel hierarchical fine-grained four-layer annotation approach. The top layer separates the posts into hateful and non-hateful categories. The second layer further categorises hateful posts into explicit hateful and implicit hateful. The third layer is the multilabel tagging of the post into topics, such as political, religion, racism, or sexism. The fourth layer involves the identification of the targeted named entity, either explicitly or implicitly. Additionally, a thorough evaluation of the data annotation schema for trustworthy annotation is provided. The HHSD data is the largest multi-layer annotated corpora in Hindi compared with the existing multi-layer annotated data. Experiments on the dataset using the transformer-based approaches in single-task learning (STL) attain encouraging performances in accuracy and weighted-f1 score. The experiment leveraged multi-task learning (MTL) by including multiple related hate speech detection tasks from high-resource English and languages from the same linguistic family such as Urdu and Bangla with a transformer encoder as the shared layers to obtain a significant increment of 5.31% and 5.35% over STL in accuracy and weighted-f1 for layer A, 8.20%, and 22.83% for layer B. The MTL surpasses STL by 8.98% and 4.07% in exact match and hamming loss for layer C.
Keyword:
Earth
Annotations
Social networking (online)
Hate speech
Tagging
Linguistics
Transformers
multi-task learning
F1 score
accuracy
Shared layers

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

I
indian institute of technology system (iit system)
学者数:
9.5W
论文数: 9.9W
被引数: 93
I
indian institute of technology (iit) - patna
学者数:
1.8K
论文数: 1.6K
被引数: 0
引用论文

引用论文

Political Economy of Climate Change Policy
err2013-01-01
err0
PREAI
errFranklin Steves; Alexander Teytelboym
err分享
err收藏
TCR density in early iNKT cell precursors regulates agonist selection and subset differentiation in mice
err2019-04-02
err0
errOAAI
errClaudine Joseph; Jihene Klibi; Ludivine Amable; Lorenzo Comba; Alessandro Cascioferro; Marc Delord; Veronique Parietti; Christelle Lenoir; Sylvain Latour; Bruno Lucas; Christophe Viret; Antoine Toubert; Kamel Benlagha
err分享
err收藏
学者 查看更多内容