Return
Discriminative semantic learning for incomplete multimodal sentiment analysis
T
H
Y
Z
DOI:10.1007/s00530-026-02555-9.png)
Abstract
En 中文
Multimodal sentiment analysis aims to recognize human affective states by integrating language, audio, and visual modalities. However, in practical deployment, sensor malfunction, transmission failures, or privacy constraints often lead to partial modality absence. Existing methods lack explicit sentiment-level semantic supervision, making it difficult for completed features to achieve sufficient sentiment discriminability. Meanwhile, gradient interference in joint optimization suppresses parameter updates of unimodal encoders, further undermining discriminative representation learning. To jointly address these two issues, we propose the Discriminative Semantic Learning framework (DiSL). The framework first introduces learnable sentiment prototypes as semantic anchors to provide explicit sentiment-discriminative guidance for feature completion. Building upon this, a gradient decoupling strategy is designed to separate the optimization paths of unimodal and multimodal objectives, preventing fusion gradients from interfering with unimodal encoders, thereby synergistically enhancing both discriminative representation learning and multimodal fusion. Extensive experiments on three benchmark datasets, CMU-MOSEI, CHERMA, and IEMOCAP, demonstrate that DiSL achieves state-of-the-art performance across various missing modality scenarios, with consistent improvements in accuracy and F1-score. Code is available at https://github.com/stao03/DiSL .
Keywords:
Multimodal sentiment analysis
Missing modality
Discriminative semantic learning
Sentiment prototype supervision
Gradient decoupling
Journal
IF:
3.1
Papers:
2.7K
Citations:
2.7K
