arrow
Return

Improving multimodal sentiment prediction through vision-language feature interaction

delete2025-01-10
delete0
PRE
AI
J
Jieyu An
B
B. Ding *
W
Wan Mohd Nazmee Wan Zainon
DOI:10.1007/s00530-024-01659-4delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Multimodal sentiment analysis aims to accurately assess the sentiment expressed in a given data source by integrating and analyzing multiple modalities, such as text and images. Extracting discriminative features for sentiment prediction is a powerful approach to address the challenges in multimodal sentiment analysis. Most methods in this domain leverage pre-trained unimodal models to extract features from individual modalities. Subsequently, these features undergo integration via sophisticated fusion mechanisms. However, these models often need to be improved in their ability to proficiently process multimodal data, potentially risking the loss of semantic associations between the different modalities. This study aims to address this problem in multimodal sentiment analysis by developing a simple end-to-end model that avoids the need for sophisticated ensemble techniques for feature extraction. In contrast, the proposed methodology capitalizes on the benefits of transfer learning through the deployment of a vision-language pre-trained model. This model efficiently extracts both visual and textual features within a cohesive framework. Extracted features are subsequently integrated via the proposed feature interaction module, which facilitates capturing potential semantic information in an image-text pair through explicit and implicit feature interaction. Finally, the derived representations undergo transmission to the classification module, thereby augmenting performance in sentiment analysis tasks. The effectiveness of the proposed approach is substantiated through a rigorous experimental evaluation. Assessments conducted on two publicly available real-world datasets reveal significant enhancements in sentiment analysis performance.
Keywords:
Multimodal deep learning
Multimodal fusion
Sentiment classification
Sentiment analysis
Vision-language pre-trained model

Journal

Multimedia Systems cover
Multimedia Systems
IF:
3.1
Papers:
2.7K
Citations:
2.7K

Organization

G
gandong university
Scholars:
37
Papers: 24
Citations: 0
U
Universiti Sains Malaysia
Scholars:
1.5W
Papers: 1.3W
Citations: 131