arrow
Return

Progressive Learning Framework With Missing Modality Reconstruction for Multimodal Emotion Recognition

delete2026-01-01
delete0
PRE
AI
H
Haoqin Sun
X
Xugang Lu
J
Jingguang Tian
J
Jiaming Zhou
H
He, Jiabei
H
Hui Wang
X
Xiangyu Kong
X
Xinhui Hu
Y
Yong Qin *
DOI:10.1109/TASLPRO.2025.3640882delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Multimodal emotion recognition (MMER) is crucial for effective human-computer interaction, yet its performance degrades sharply when modality data are incomplete due to privacy constraints or data loss. Existing approaches often attempt to reconstruct missing information directly from raw, heterogeneous modalities, requiring complex cross-modal mappings that lead to low-quality reconstruction and hence less performance improvement. To address these challenges, we propose a progressive learning framework (ProLF) for missing-modality reconstruction and robust multimodal fusion. ProLF's core innovation lies in its design of an incremental three-stage strategy, i.e., Alignment, Reconstruction, and Augmentation-which collectively mitigates modality heterogeneity, data incompleteness, and weak fused representations. In the Alignment stage, ProLF applies unsupervised distribution-level contrastive learning to reduce cross-modal discrepancies and establish a unified latent space. In the Reconstruction stage, a normalizing flow model efficiently transforms the aligned latent distributions to recover missing modalities with high accuracy. In the Augmentation stage, a mixture of modality experts performs deep fusion of the restored multimodal information, while supervised point-based contrastive learning enhances emotion-specific features and suppresses irrelevant semantics, substantially improving discriminability. Extensive experiments on multiple benchmark datasets validate the effectiveness of ProLF. On IEMOCAP and MSP-IMPROV, ProLF achieves average performance gains of 1.5%-2%, and on CMU-MOSI, improvements of 1%-3%. These results confirm that ProLF provides robust and consistent benefits under both missing and complete modality conditions.
Keywords:
Emotion recognition
Semantics
Feature extraction
Contrastive learning
Image reconstruction
Computer architecture
Training
Visualization
Transforms
Speech processing
Multimodal emotion recognition
mixture of modality expert
missing modality

Journal

I
IEEE Transactions on Audio Speech and Language Processing
IF:
0
Papers:
151
Citations:
0

Organization

U
university of exeter
Scholars:
2.6K
Papers: 1.4K
Citations: 0
N
nankai university
Scholars:
4.7W
Papers: 3.2W
Citations: 74
researcher View more organizations