Return
Collaborative fine-grained interaction learning for image-text sentiment analysis
DOI:10.1016/j.knosys.2023.110951.png)
Abstract
En 中文
Investigating interactions between image and text can effectively improve image-text sentiment analysis, but most existing methods do not explore image-text interaction at fine-grained level. In this paper, we propose a Memory-enhanced Collaborative Fine-grained Interaction Transformer (MCFIT) to learn collaborative fine-grained interaction between image and text. Specifically, a multi-branch encoder is designed to learn both fine-grained region-word and patch-word interactions. Meanwhile, Memory-enhanced Cross-Attention (MECA) is proposed to utilize patch and region information to improve region-word interaction and patch-word interaction learning, respectively. Therefore, collaborative fine-grained interaction can yield more accurate image-text interaction. Finally, to analyze the sentiments embedded in real-life Chinese image-text pairs, we build a large-scale Chinese image-text sentiment dataset (CISD) containing 54,931 image-text pairs. Extensive experiments conducted on four real-life datasets prove the effectiveness of collaborative fine-grained interaction and demonstrate that MCFIT outperforms the state-of-the-art baselines.(c) 2023 Elsevier B.V. All rights reserved.
Keywords:
Image-text sentiment analysis
Fine-grained interaction
Image-text dataset
Memory transformer

