1
Return

Tri-cross modal fusion network for multimodal sentiment analysis in short videos

delete2026-05-01
delete0
PRE
AI
H
Heyong Wang *
Y
Yuanhao Chen
DOI:10.1016/j.image.2026.117542delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
With the shift of online social entertainment from text to short videos, user-generated content has become increasingly multimodal, conveying emotional information through text, vision, and audio. Traditional unimodal sentiment methods struggle with such complexity: facial expressions shift meaning with context; sarcasm and irony rely on prosody and contextual cues; and discrepancies between spoken words and facial displays are common, making single-modality approaches inadequate for modeling authentic affect. To address these challenges, this paper proposes a Tri-Cross Modal Fusion Network (TCMFN) for multimodal sentiment analysis. During the feature extraction phase, for the visual modality, visual features are extracted from video frames using the pretrained Vision Transformer (ViT), and combined with Facial Action Units (FAU) to enhance sensitivity to subtle emotional variations. For the textual modality, an Emotion CLIP-based encoder is utilized in conjunction with emotion description templates, jointly encoding video frames and template sentences, selecting the most appropriate descriptive templates, and concatenating them with the video transcript to enrich contextual representation. For the audio modality, the Data2Vec model is employed to extract deep acoustic features. During multimodal fusion, a Tri-modal Cross-attention Encoder (TCME) is designed to facilitate inter-modal interactions, and at the fusion stage's end, a Multi-head Feature Weighting Attention Fusion (MFWAF) mechanism dynamically optimizes feature weights from different modalities. Experimental results on the CMU-MOSI and CMU-MOSEI datasets demonstrate that TCMFN outperforms current mainstream methods, validating its effectiveness and performance.
Keywords:
Multimodal sentiment analysis
Multimodal feature enhancement
Cross-modal attention mechanism
Multimodal data fusion

Journal

S
SIGNAL PROCESSING-IMAGE COMMUNICATION
IF:
2.7
Papers:
18
Citations:
0

Organization

S
south china university of technology
Scholars:
6.6W
Papers: 5.0W
Citations: 85
Cited Papers

Cited Papers

Citing Papers

Citing Papers