arrow
Return

Multimodal sentiment analysis method based on image-text quantum transformer

delete2025-07-01
delete0
PRE
AI
H
Huaiguang Wu
D
Delong Kong
L
Lijie Wang
D
Daiyi Li *
J
Jiahui Zhang
Y
Yucan Han
DOI:10.1016/j.neucom.2025.130107delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Multimodal sentiment analysis aims to recognize and interpret the diverse emotional information contained in different modal data. Although multimodal sentiment analysis methods have achieved significant results, they still have certain limitations in capturing the complex features of high-dimensional data. This paper proposes a multimodal sentiment analysis method utilizing the Image and Text Quantum Transformer (ITQT-MSA), which innovatively combines quantum computing and deep learning technologies to achieve more accurate emotion recognition and analysis. Firstly, a text-oriented parametric quantum circuit is designed that exploits quantum superposition and entanglement properties and combines the transformer model to achieve deep feature extraction from text data. Secondly, another image-oriented parametric quantum circuit is constructed and combined with the Visual Transformer (VIT) to adequately extract the emotional information in the image. Finally, effective alignment and the integration of features from both text and images are achieved by the designed cross-attention fusion mechanism. The experiments are performed on classical computers by designing quantum circuits within a simulated noisy environment, and the results show that the proposed method outperforms SOTA models.
Keywords:
Multimodal sentiment analysis
Quantum transformer
Parameterized quantum circuit
Visual Transformer
Cross-attention fusion

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

No organization information available