Return
Reconstructing representations using diffusion models for multimodal sentiment analysis through reading comprehension
DOI:10.1016/j.asoc.2024.112346.png)
Abstract
En 中文
The primary challenge in multimodal sentiment analysis (MSA), which utilizes textual, audio, and visual information to analyze speakers' emotions, lies in constructing representation vectors that incorporate both unimodal semantic and multimodal interaction information. While existing research has extensively focused on multimodal fusion strategies, there remains insufficient exploration of the intrinsic potential within concurrently enhancing unimodal and multimodal representations. To address this gap, we introduce two additional steps to traditional three-step MSA: text modality enhancement through the machine reading comprehension (MRC) framework and multimodal representation reconstruction via proposing the diverse diffusion denoising autoencoder (D3AE) module. The MRC queries are integrated to locate sentiment-related prior knowledge, thereby deepening textual semantic understanding from a pretrained language model. Meanwhile, D3AE employs a single-step denoising strategy along with diffusion models across multiple time intervals, enabling efficient reconstruction and enhancement of multimodal representations. Extensive experiments conducted on two benchmark datasets, CMU-MOSI and CMU-MOSEI, validate that our model, MRC-D3AE, achieves state-of-the-art performance. The superiority of our model over existing baselines is primarily attributed to integrating MRC for enhancing text modality and D3AE for reconstructing multimodal representations.
Keywords:
Multimodal sentiment analysis
Machine reading comprehension
Diffusion model
Diffusion denoising autoencoder
Multimodal fusion representation
Journal
IF:
6.6
Papers:
1.4W
Citations:
4.8W
Organization
Cited Papers
ReCoMIF: Reading comprehension based multi-source information fusion network for Chinese spoken language understanding
INFORMATION FUSION
IF15.5

