Return
Multi-view dynamic perception framework for Chinese harmful meme detection
DOI:10.1016/j.ipm.2025.104602.png)
Abstract
En 中文
Chinese harmful memes convey toxicity on social media through varied semantic complexity and diverse modality combinations. However, existing detection methods typically adopt static architectures with fixed interaction patterns, which lack the flexibility to accurately identify harmful cues embedded in heterogeneous semantic and modal content across different memes, hindering a comprehensive understanding of toxic intent. To address this limitation, we propose the Multi-View Dynamic Perception (MDP) framework, a dynamic interaction paradigm specifically designed for Chinese harmful meme detection. Specifically, we develop five types of semantic perception nodes to synchronously extract features from diverse views. These nodes are densely stacked to form two perception branches, respectively guided by textual and visual features, to effectively capture modality-specific cues. To enhance adaptability, each node is equipped with an independent soft router that dynamically regulates information flow and enables flexible interaction patterns tailored to different memes. Furthermore, we introduce a Hierarchical Mutual Learning module to promote complementary representation learning between the two branches via mutual information maximization. Extensive experiments on the publicly available dataset TOXICN MM, comprising 12,000 samples, demonstrate the effectiveness of the proposed framework, with F1 score improvements of 1.06% in harmful meme detection and 2.77% in harmful type identification over the previous state-of-the-art method. We further evaluate the generalization of the MDP framework on a Chinese multimodal sarcasm detection dataset, where the proposed method also achieves competitive results.
Journal
I
IF:
6.9
Papers:
323
Citations:
0

