Return
Multi-modal sarcasm detection with dynamic modality fusion and multi-frequency information propagation
Y
J
L
L
J
DOI:10.1007/s00530-026-02526-0.png)
Abstract
En 中文
Sarcasm is a linguistic phenomenon in which the expression contradicts the true intention, highlighting the inconsistency between the literal meaning and the intended meaning. Mining the importance of different modalities and the inconsistent information between modalities is a key challenge in recognizing multi-modal sarcasm. In recent years, significant progress has been made in modeling image-text inconsistencies using attention mechanisms and graph convolutional networks. However, these methods have some limitations. On the one hand, they rely on static networks that cannot flexibly assess the importance of different modalities when handling samples with varying modality informativeness, such as text-dominant, image-dominant, or jointly informative cases. On the other hand, graph convolutional networks tend to ignore high-frequency information, which is an important means of reflecting inconsistencies between nodes. To address the above issues, we designed a dynamic fusion network dynamically estimate the contribution of each modality for different sample types and adaptively fuse text and image information. Additionally, we applied a graph convolutional network based on multi-frequency information filters, which can adaptively integrate low-frequency and high-frequency information between nodes during the information aggregation phase, thus better learning node representations. Finally, we also applied contrastive learning to further differentiate the feature representations of sarcastic and non-sarcastic samples. We conducted experiments on two established multimodal sarcasm detection benchmarks, both of which achieved state-of-the-art performance.
Keywords:
Multi-modal sarcasm detection
Modality fusion
Multi-frequency information propagation
Journal
IF:
3.1
Papers:
2.7K
Citations:
2.7K
