Return
Emotionally Controllable Audio-driven Talking Face Generation
DOI:10.1145/3779219.png)
Abstract
En 中文
Talking face generation is a technique that synthesizes realistic, speech-synchronized facial animations from static images driven by audio or textual input. In recent years, notable advancements have been achieved in the generation of realistic lip-synchronized talking face videos. Although these methods have demonstrated highly realistic results in lip-sync generation, they rarely focus on the expression of emotions, which is essential for achieving emotionally rich and lifelike talking face videos. In this article, we propose a novel emotionally controllable audio-driven talking face generation framework, termed ECATFG. Specifically, the proposed ECATFG mainly consists of three modules: template video generator, landmark generator, and rendering module. Firstly, we integrate a customized motion transfer module into StyleGAN to generate template videos that encapsulate the target emotions and head movements. Subsequently, a Mamba-based landmark generator is proposed to accurately model the correlation between input audio and facial landmarks. Afterwards, in the rendering module, we employ a framework that combines both AdaIN and SPADE layers to transform audio-predicted landmarks into highly realistic images, guided by reference images and facial contour landmarks. Finally, comprehensive experiments validate the effectiveness of ECATFG, demonstrating its capability to generate high-quality talking face videos that not only achieve exceptional lip-synchronization accuracy but also effectively convey a wide range of emotions.
Keywords:
Talking face generation
lip synchronization
emotion
face editing
Journal
IF:
6
Papers:
2.0K
Citations:
5.4K

