arrow
Return

Reconstructing representations using diffusion models for multimodal sentiment analysis through reading comprehension

delete2024-12-01
delete0
PRE
AI
H
Hua Zhang *
Z
Zijing Cai
陈
陈碧 (Bi Yu Chen)
B
Bo Jiang
B
Bo Xie
DOI:10.1016/j.asoc.2024.112346delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The primary challenge in multimodal sentiment analysis (MSA), which utilizes textual, audio, and visual information to analyze speakers' emotions, lies in constructing representation vectors that incorporate both unimodal semantic and multimodal interaction information. While existing research has extensively focused on multimodal fusion strategies, there remains insufficient exploration of the intrinsic potential within concurrently enhancing unimodal and multimodal representations. To address this gap, we introduce two additional steps to traditional three-step MSA: text modality enhancement through the machine reading comprehension (MRC) framework and multimodal representation reconstruction via proposing the diverse diffusion denoising autoencoder (D3AE) module. The MRC queries are integrated to locate sentiment-related prior knowledge, thereby deepening textual semantic understanding from a pretrained language model. Meanwhile, D3AE employs a single-step denoising strategy along with diffusion models across multiple time intervals, enabling efficient reconstruction and enhancement of multimodal representations. Extensive experiments conducted on two benchmark datasets, CMU-MOSI and CMU-MOSEI, validate that our model, MRC-D3AE, achieves state-of-the-art performance. The superiority of our model over existing baselines is primarily attributed to integrating MRC for enhancing text modality and D3AE for reconstructing multimodal representations.
Keywords:
Multimodal sentiment analysis
Machine reading comprehension
Diffusion model
Diffusion denoising autoencoder
Multimodal fusion representation

Journal

Applied Soft Computing cover
Applied Soft Computing
IF:
6.6
Papers:
1.4W
Citations:
4.8W

Organization

Z
Zhejiang Gongshang University
Scholars:
6.6K
Papers: 4.9K
Citations: 8.1K
Cited Papers

Cited Papers

Fairness in cost-benefit analysis: A methodology for health technology assessment
err2017-06-16
err0
PREAI
errAnne-Laure Samson; Erik Schokkaert; Clémence Thébaut; Brigitte Dormont; Marc Fleurbaey; Stéphane Luchini; Carine Van de Voorde
errShare
errSave
A Quantum Algorithm for System Specifications Verification
err2024-07-15
err3
errOAAI
errZidan, Mohammed; Eisa, Ahmed M.; Qasymeh, Montasir; Shoman, Mahmoud A. Ismail
errShare
errSave
Identification of rainfall homogenous regions in Saudi Arabia for experimenting and improving trend detection techniques
err2021-11-27
err0
PREAI
errJaved Mallick; Swapan Talukdar; Mohammed K. Almesfer; Majed Alsubih; Mohd. Ahmed; Abu Reza Md. Towfiqul Islam
errShare
errSave
errShare
errSave
StyleBERT: Text-audio sentiment analysis with Bi-directional Style Enhancement
err2023-03-01
err7
PREAI
errLin, Fei; Liu, Shengqiang; Zhang, Cong; Fan, Jin; Wu, Zizhao
errShare
errSave
ReCoMIF: Reading comprehension based multi-source information fusion network for Chinese spoken language understanding
err2023-08-01
err8
PREAI
errXie, Bo; Jia, Xiaohui; Song, Xiawen; Zhang, Hua; Chen, Bi; Jiang, Bo; Wang, Ye; Pan, Yun
errShare
errSave
Diffusion Models: A Comprehensive Survey of Methods and Applications
err2023-11-09
err245
errOAAI
errYang, Ling; Zhang, Zhilong; Song, Yang; Hong, Shenda; Xu, Runsheng; Zhao, Yue; Zhang, Wentao; Cui, Bin; Yang, Ming-Hsuan
errShare
errSave
Targeted Aspect-Based Multimodal Sentiment Analysis: An Attention Capsule Extraction and Multi-Head Fusion Network
err2021-01-01
err28
errOAAI
errGu, Donghong; Wang, Jiaqian; Cai, Shaohua; Yang, Chi; Song, Zhengxin; Zhao, Haoliang; Xiao, Luwei; Wang, Hua
errShare
errSave
researcher View more