arrow
Return

Modality emotion semantic correlation analysis for multimodal emotion recognition

delete2025-06-10
delete0
PRE
AI
Y
Yuqing Zhang
D
Dongliang Xie
D
Dawei Luo
B
Baosheng Sun
DOI:10.1016/j.compeleceng.2025.110467delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Affective computing serves as the fundamental technology and a crucial prerequisite for attaining naturalized and anthropomorphic human–computer interaction. Nevertheless, the expression of emotion is complex and multi-dimensional, posing significant challenges for multimodal emotion recognition due to the heterogeneity gap among distinct modalities. To tackle this issue, we propose a novel approach named modality emotion semantic correlation analysis (MESCA), which enhances multimodal affective semantic consistency by leveraging modality correlation learning to achieve multimodal information complementation. Specifically, we first design a modal-pair correlation module that calculates emotion semantic consistency across text, audio and video information. This module contributes to a comprehensive understanding of the emotional state by fusing complementary semantic information and assists in mitigating redundancy in pairwise interaction methods. Next, we introduce structural re-parameterization technology that transforms the multi-branch training structure into a single-branch inference structure to solve the problem of excessive computational expense, thereby facilitating a more efficient and effective recognition process. Additionally, the proposed model is verified on two public datasets, IEMOCAP and CMU-MOSEI. Compared to baseline methods, MESCA significantly enhances efficiency while maintaining prediction accuracy on IEMOCAP, and outperforms on both efficiency and accuracy on CMU-MOSEI.

Journal

C
Computers and Electrical Engineering
IF:
4.9
Papers:
6.7K
Citations:
1.3W

Organization

No organization information available