arrow
Return

Multi-task learning framework using tri-encoder with caption prompt for multimodal aspect-based sentiment analysis

delete2025-04-29
delete0
PRE
AI
Y
Yuanyuan Cai
童
童飞 (Fei Tong)
张
张青川 (Qingchuan Zhang) *
H
Haitao Xiong
DOI:10.1007/s11227-025-07252-4delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Multimodal aspect-based sentiment analysis (MABSA) is an advanced technology to identify all aspect terms together with their respective sentiment mentioned in the multimodal data. Most existing MABSA methods encounter two main challenges: (1) The representations used in MABSA are directly encoded by the general pre-trained models, which are insensitive to identifying aspect-level sentiment. (2) Latent visual semantic information is underutilized when representing the key aspects and their sentimental polarities. To address the mentioned challenges, we propose an optimized multi-task learning framework using a tri-encoder with caption prompt (TECP), including MABSA and two auxiliary unimodal tasks to jointly learn the aspect-aware and sentiment-aware multimodal representation. In TECP, the merged-attention fusion network is designed to obtain multimodal features, which enhances the intra-modal and cross-modal semantic interactions. Within the tri-encoder, the caption encoder is designed to further generate visual caption feature as important clue to enrich the multimodal semantic and sentimental information. Moreover, within the caption encoder, the dependency weight attention network is proposed to focus on the aspect-level feature in the caption sentence. We conduct elaborate experiments and evaluate the performance of TECP with respect to Precision, Recall, and F1-score. Our TECP achieves SOTA results on two benchmark Twitter datasets in comparison with previous baseline models.
Keywords:
Multimodal aspect-based sentiment analysis
Multi-task learning
Image caption prompt
Dependency weight attention
Merged-attention fusion network

Journal

Journal of Supercomputing cover
Journal of Supercomputing
IF:
2.7
Papers:
1.1K
Citations:
1.0W

Organization

B
Beijing Technology and Business University
Scholars:
4.1K
Papers: 1.7K
Citations: 1.6W
Cited Papers

Cited Papers

Bi-Bimodal Modality Fusion for Correlation-Controlled Multimodal Sentiment Analysis
err2021-10-18
err0
errOAAI
errWei Han; Hui Chen; Alexander Gelbukh; Amir Zadeh; Louis-philippe Morency; Soujanya Poria
errShare
errSave
Atlantis: Aesthetic-oriented multiple granularities fusion network for joint multimodal aspect-based sentiment analysis
err2024-06-01
err12
PREAI
errXiao, Luwei; Wu, Xingjiao; Xu, Junjie; Li, Weijie; Jin, Cheng; He, Liang
errShare
errSave
errShare
errSave
errShare
errSave
researcher View more