返回
Learning coordinated emotion representation between voice and face
DOI:10.1007/s10489-022-04216-6.png)
摘要
En 中文
Voice and face information are two most important perceptual modalities for human. In recent years, many researchers show great interest in learning cross-modal representations for different face-voice association tasks. However, these existing methods focus on the various biological attributions but rarely take emotion semantics between voice and face into account. In this paper, we present a novel two-stream model, called Emotion Representation Learning Network (EmoRL-Net), to learn the cross-modal coordinated emotion representations for various downstream matching and retrieval tasks. Within the proposed approach, we first propose two sub-network architectures that learn two unimodal features from the two modalities. Afterwards, we train EmoRL-Net by an objective loss function which includes one explicit and two implicit constraints. Meanwhile, an online semi-hard negative mining strategy is utilized to construct triplet units in a mini-batch manner, thereby stabilize and speeding up the learning process. Extensive experiments demonstrate that the proposed method can benefit various face-voice emotion tasks, including cross-modal verification, 1:2 matching, 1:N matching, and retrieval scenarios. The experiment results also show the proposed method outperforms the state-of-the-art approaches.
Keyword:
Face-voice emotion relationship
Coordinated emotion representation
Cross-modal matching
Metric learning
期刊
IF:
3.5
论文数:
7.6K
被引数:
1.7W
机构
引用论文
A review of affective computing: From unimodal analysis to multimodal fusion情感计算综述: 从单峰分析到多模态融合
INFORMATION FUSION
IF15.5

