Return
Multimodal sentiment analysis method based on image-text quantum transformer
DOI:10.1016/j.neucom.2025.130107.png)
Abstract
En 中文
Multimodal sentiment analysis aims to recognize and interpret the diverse emotional information contained in different modal data. Although multimodal sentiment analysis methods have achieved significant results, they still have certain limitations in capturing the complex features of high-dimensional data. This paper proposes a multimodal sentiment analysis method utilizing the Image and Text Quantum Transformer (ITQT-MSA), which innovatively combines quantum computing and deep learning technologies to achieve more accurate emotion recognition and analysis. Firstly, a text-oriented parametric quantum circuit is designed that exploits quantum superposition and entanglement properties and combines the transformer model to achieve deep feature extraction from text data. Secondly, another image-oriented parametric quantum circuit is constructed and combined with the Visual Transformer (VIT) to adequately extract the emotional information in the image. Finally, effective alignment and the integration of features from both text and images are achieved by the designed cross-attention fusion mechanism. The experiments are performed on classical computers by designing quantum circuits within a simulated noisy environment, and the results show that the proposed method outperforms SOTA models.
Keywords:
Multimodal sentiment analysis
Quantum transformer
Parameterized quantum circuit
Visual Transformer
Cross-attention fusion
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W
Organization
No organization information available

