1
Return

PoseTalk: Exploring Text- and Audio-Based Pose Control for One-Shot Talking Face Generation

delete2026-03-01
delete0
PRE
AI
J
Jun Ling *
W
Wang, Yiwen
H
Han Xue
S
Song, Li
DOI:10.1145/3788691delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Although audio-driven talking face generation has witnessed significant advances in recent years, two problems remain to be solved. First, the existing models cannot control the long-term head actions as humans expect, because the audio can only provide short-term cues, such as the rhythms and sentiments, for head movements. Second, generating long-term head poses and ensuring accurate lip motions remain challenging due to the difficulty in harnessing the optimization process for large-scale head movements and small-scale mouth motions. In this study, we propose a novel method to address these issues. First, to alleviate the limitations of audio conditions, we propose a Pose Latent Diffusion (PLD) model to generate head motions from two kinds of input modalities: the input audio and user-controlled text prompts. The audio provides short-term rhythm correspondence with the head movements, while the text prompts describe the long-term semantics of head motions. Second, we propose a refinement-based learning strategy to synthesize head movements and accurate lip motions using two cascaded networks, namely CoarseNet and RefineNet. The CoarseNet estimates coarse global motions to produce animated images with changed poses, and the RefineNet progressively estimates finer lip motions from low to high resolutions, yielding improved lip-synchronization performance. Experiments demonstrate that our method can achieve better pose diversity and realness compared to audio-based pose generation baselines, and our video generator model outperforms state-of-the-art methods in synthesizing natural head motions. Projects and demos are available at https://junleen.github.io/projects/posetalk.
Keywords:
Talking face generation
audio-driven
pose control

Journal

ACM Transactions on Multimedia Computing Communications and Applications cover
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
Papers:
2.0K
Citations:
5.4K

Organization

S
shanghai jiao tong university
Scholars:
15.1W
Papers: 11.5W
Citations: 159
Cited Papers

Cited Papers

Citing Papers

Citing Papers