arrow
Return

VAE-Integrated Multiscale Generative Diffusion Modeling for Incomplete Multimodal Emotion Recognition

delete2026-04-28
delete0
PRE
AI
Z
ZiXuan Wang
Z
ZhuYi Yao
J
JiaYue Shen
P
Pan Wang
X
Xuejiao Chen
张旭 cover
张旭 (Xu Zhang)
H
HuangLiang Gu
X
Xiaokang Zhou
DOI:10.1109/TCSS.2026.3681252delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Multimodal emotion recognition (MER) has been widely adopted in affective computing and human–computer interaction. However, real-world multimodal streams frequently suffer from missing or incomplete modalities due to sensor failures, privacy constraints, and heterogeneous acquisition costs, which often causes severe performance degradation for models trained under complete inputs. To address this issue, we propose hierarchical VAE–diffusion for emotion reconstruction (HVDER), a latent-space generative framework for incomplete MER. HVDER first employs modality-specific variational autoencoders (VAEs) to project language, visual, and audio features into a unified low-dimensional latent space with KL-regularized structure. Then, a two-stage coarse-to-fine conditional diffusion module completes missing modality latents by recovering global emotion semantics followed by refining local discriminative details. Finally, the completed and observed modality representations are fused for emotion prediction, optimized end-to-end with a multiobjective loss integrating reconstruction, diffusion matching, cross-modal alignment, and classification supervision. Extensive experiments on CMU-MOSEI and CMU-MOSI validate the effectiveness and robustness of HVDER. Under fixed modality-missing settings, HVDER achieves average ACC2/F1 of 76.8/75.4 on CMU-MOSEI and 73.4/72.9 on CMU-MOSI, outperforming representative baselines across all modality combinations. Under random missing with a high missing rate of MR = 0.7, HVDER maintains ACC2/F1 of 75.4/72.5 on CMU-MOSEI and 68.0/67.2 on CMU-MOSI, indicating a stronger performance lower bound in severely incomplete scenarios. Ablation studies further confirm that latent-space modeling and the coarse-to-fine diffusion design jointly contribute to the main performance gains.
Keywords:
Emotion AI
fusion reconstruction
multimodal emotion recognition (MER)

Journal

IEEE Transactions on Computational Social Systems cover
IEEE Transactions on Computational Social Systems
IF:
4.9
Papers:
577
Citations:
6.8K

Organization

B
bnpp consumer finance company ltd.
Scholars:
2
Papers: 2
Citations: 0
U
university of post & telecommunications
Scholars:
8
Papers: 2
Citations: 0
U
university of east anglia
Scholars:
1.1K
Papers: 605
Citations: 0
K
Kansai University
Scholars:
1.7K
Papers: 1.5K
Citations: 0
V
vocational college of information technology
Scholars:
2
Papers: 2
Citations: 0
researcher View more organizations