1
Return

Precise Speech-Driven Talking-Face Synthesis with Realistic Speaker-Emulated Facial Expressions

delete2026-04-01
delete0
PRE
AI
W
Wang, Jhing-Fa
L
Liou, Shu-Yan *
T
Tsai, Hsin-Chun
DOI:10.1142/S0218001426570089delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Traditional facial animation models do not particularly stress the naturalness of imitating the dynamic facial expressions, including eye movements, dynamic eyebrow deformation, and head motion naturalness. Hence, in this paper, a speech-driven talking-face synthesizer (SDTS) is proposed for generating the dynamic talking video of a given static face for semantically mimicking the speech of any real person. The SDTS can lead the static digital-twin face to vividly mimic the expressive motions of the face and lip-synced mouth of various speakers with a personalized accent with high distinctiveness. The SDTS framework has two stages. In the first stage, one branch, termed the dynamic fused-features generation module (DFGM), contains a cross-modal speech-facial fusion module (CSFF) and a temporal convolutional network (TCN). The CSFF is the core to seamlessly align the speech features and facial features. The second branch is the self-designed adaptive identity extractor (AIE), where a series of residual blocks using partial batch normalization unit (PBN-ResNet blocks) and the residual blocks with the squeeze-and-excitation unit (SE-ResNet blocks) are cascaded to precisely capture the key features of the face in a static reference image. In the second stage of SDTS, the diffusion model termed diffusion-based rendering model (DIRM) is applied to generate the high-resolution video reconstruction with the consistencies of appearance and emotion via fusing the driving speech features and the referred facial features. The extensive experiments demonstrate that SDTS can significantly promote lip-synchronization, enrich the upper facial expression, and exhibit the naturalness of the head movements. Moreover, the SDTS can steadily maintain the facial identity consistency and the facial expression coherence for varying speaking speeds and emotions. Hence, it can attain less than 5.26 FID, 0.72 LSE-D, and 0.56 LME than the StyleTalk model, which is a well-known talking-face synthesis model.
Keywords:
Speech-driven talking face generation
lip synchronization
diffusion-based generation

Journal

International Journal of Pattern Recognition and Artificial Intelligence cover
International Journal of Pattern Recognition and Artificial Intelligence
IF:
1.1
Papers:
161
Citations:
2.0K

Organization

National Chiayi University cover
National Chiayi University
Scholars:
1.9K
Papers: 2.0K
Citations: 1.4K
N
national cheng kung university
Scholars:
3.0K
Papers: 1.2K
Citations: 0
Sanda University cover
Sanda University
Scholars:
111
Papers: 97
Citations: 125
C
Cheng Shiu University
Scholars:
511
Papers: 713
Citations: 506
Cited Papers

Cited Papers

Citing Papers

Citing Papers