arrow
Return

Collaborative fine-grained interaction learning for image-text sentiment analysis

delete2023-11-01
delete3
PRE
AI
普园媛 cover
普园媛 (Yuanyuan Pu) *
D
Dongming Zhou
曹进德 (Jinde Cao)
J
Jinjing Gu
Z
Zhengpeng Zhao
D
Dan Xu
DOI:10.1016/j.knosys.2023.110951delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Investigating interactions between image and text can effectively improve image-text sentiment analysis, but most existing methods do not explore image-text interaction at fine-grained level. In this paper, we propose a Memory-enhanced Collaborative Fine-grained Interaction Transformer (MCFIT) to learn collaborative fine-grained interaction between image and text. Specifically, a multi-branch encoder is designed to learn both fine-grained region-word and patch-word interactions. Meanwhile, Memory-enhanced Cross-Attention (MECA) is proposed to utilize patch and region information to improve region-word interaction and patch-word interaction learning, respectively. Therefore, collaborative fine-grained interaction can yield more accurate image-text interaction. Finally, to analyze the sentiments embedded in real-life Chinese image-text pairs, we build a large-scale Chinese image-text sentiment dataset (CISD) containing 54,931 image-text pairs. Extensive experiments conducted on four real-life datasets prove the effectiveness of collaborative fine-grained interaction and demonstrate that MCFIT outperforms the state-of-the-art baselines.(c) 2023 Elsevier B.V. All rights reserved.
Keywords:
Image-text sentiment analysis
Fine-grained interaction
Image-text dataset
Memory transformer

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

Y
Yunnan University
Scholars:
1.6W
Papers: 9.9K
Citations: 13