Return
MPF: A Multi-Level Perceiving Framework for Multimodal Sarcasm Detection
DOI:10.1016/j.inffus.2026.104436.png)
Abstract
En 中文
Multimodal sarcasm detection aims to identify sarcastic expressions by integrating information from multiple modalities, such as images, audio, and text, while capturing emotional inconsistencies across these modalities. However, most previous approaches focus on designing static, fixed network architectures that primarily emphasize detecting cross-modal incongruities as cues for sarcasm. These methods often overlook the importance of dynamically adjusting cross-modal attention when interpreting complex and subtle sarcastic communication, especially in conversational settings. To address these limitations, we propose a Multi-level Perception Framework (MPF)—a novel approach for dynamically learning and coordinating both coarse- and fine-grained features within and across modalities for multimodal sarcasm detection. MPF consists of three core components: the Multi-View Interaction Module, the Cross-modal Routing Interaction Module, and the Cross-modal Semantic-guided Incongruity Learning Module. These components allow the framework to flexibly adjust cross-modal focus based on context while capturing the inherent emotional semantics in the fine-grained features of each modality. Extensive experiments on multiple public datasets and comparative analyses with state-of-the-art methods demonstrate the effectiveness and superiority of the proposed framework. Both qualitative and quantitative evaluations highlight MPF’s adaptability and enhanced performance in handling diverse sarcastic scenarios.
Keywords:
Multimodal sarcasm detection
Cross-modal attention
Emotional incongruity
Multi-level perception
Conversational sarcasm
Journal
IF:
15.5
Papers:
4.1K
Citations:
2.7W

