Return
Semantic Aggregated Adversarial Training Framework for Hate Speech Detection
DOI:10.1287/ijoc.2023.0508.png)
Abstract
En 中文
Hate speech poses a growing challenge to digital platforms with large and diverse user bases, prompting widespread adoption of deep learning (DL) models for automated detection at scale. Existing research, however, predominantly focuses on improving detection accuracy while paying limited attention to the vulnerability of DL-based detection models to adversarial attacks from malicious spreaders. To bridge this gap, we propose an adversarial training framework to improve the adversarial robustness of hate speech detection. This framework integrates imbalanced adversarial training with a novel semantic aggregation technology to learn robust yet discriminative features from hate speech corpora. We further introduce an adversarial attack generation framework to assess the performance of existing DL-based hate speech detection models under such attacks. Extensive computational experiments conducted on eight publicly available hate speech corpora demonstrate the robustness of the proposed method against attacks. In contrast, we show that existing DL-based detection models can be easily circumvented by adversarial attacks, allowing the dissemination of hateful sentiments through subtle modifications to the content. Additionally, we conduct comparative analyses of the proposed method with various adversarial training and imbalance training methods to illustrate its effectiveness in simultaneously addressing the data imbalance and feature inseparability issues inherent in hate speech detection. This study presents significant managerial implications, aiding online platforms in implementing effective measures to prevent the spread of hateful speech.
Keywords:
hate speech detection
adversarial attack
adversarial training
semantic aggregation
data imbalance
responsible AI
Journal
I
IF:
2.1
Papers:
86
Citations:
3.2K

