Return
Dynamic cross-instance context mining for multimodal sentiment analysis
DOI:10.1016/j.ipm.2026.104921.png)
Abstract
En 中文
Contextual modeling is pivotal in Multimodal Sentiment Analysis (MSA). Existing methods primarily rely on an instance-centric paradigm, which focuses on intra-modal dependencies (refining single-modality features) and inter-modal interactions (fusing cross-modality cues) within a single utterance. This isolated approach fails to exploit cross-instance semantic correlations, making models susceptible to local modality noise and incomplete expressions. To bridge this gap, we propose DC2M, a framework that shifts the focus from static single-sample fusion to dynamic cross-instance context mining. Our technical highlights are twofold: (i) the Adaptive Context Mining Module (ACMM), which captures cross-instance emotional evolution by retrieving the top-K relevant samples as semantic anchors, effectively increasing the contextual information gain; and (ii) the Heterogeneous Feature Fusion Module (HFFM), which optimizes intra-sample complementarity. Unlike traditional models that treat each utterance independently, DC2M performs utterance-level sentiment prediction enhanced by cross-instance contextual information from semantically related utterances. Experiments on the CMU-MOSI and CMU-MOSEI benchmarks show that DC2M achieves competitive performance, reaching 89.33% and 87.62% Acc-2, respectively. Notably, DC2M outperforms fine-tuned Qwen-1.8B by 3.31% and ChatGLM3-6B by 0.77% in Acc-7, despite a substantially more lightweight architecture. Ablation studies confirm that ACMM and HFFM synergistically enhance contextual representation learning.
Keywords:
Multimodal sentiment analysis
Sentiment classification
Modality fusion
Deep learning
Journal
I
IF:
6.9
Papers:
544
Citations:
0
Organization
Cited Papers
No cited papers available

