arrow
返回

StableFace: Analyzing and Improving Motion Stability for Talking Face Generation

delete2023-11-01
delete3
delete
OA
AI
J
Jun Ling
X
Xu Tan
L
Liyang Chen
R
Runnan Li
Y
Yuchao Zhang
S
Sheng Zhao
李
李松 (Li Song) *
DOI:10.1109/JSTSP.2023.3333552delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
While previous methods for speech-driven talking face generation have shown significant advances in improving the visual and lip-sync quality of the synthesized videos, they have paid less attention to lip motion jitters which can substantially undermine the perceived quality of talking face videos. What causes motion jitters, and how to mitigate the problem? In this article, we conduct systematic analyses to investigate the motion jittering problem based on a state-of-the-art pipeline that utilizes 3D face representations to bridge the input audio and output video, and implement several effective designs to improve motion stability. This study finds that several factors can lead to jitters in the synthesized talking face video, including jitters from the input face representations, training-inference mismatch, and a lack of dependency modeling in the generation network. Accordingly, we propose three effective solutions: 1) a Gaussian-based adaptive smoothing module to smooth the 3D face representations to eliminate jitters in the input; 2) augmented erosions added to the input data of the neural renderer in training to simulate the inference distortion to reduce mismatch; 3) an audio-fused transformer generator to model inter-frame dependency. In addition, considering there is no off-the-shelf metric that can measures motion jitters of talking face video, we devise an objective metric (Motion Stability Index, MSI) to quantitatively measure the motion jitters. Extensive experimental results show the superiority of the proposed method on motion-stable talking video generation, with superior quality to previous systems.
Keyword:
Talking face generation
vision transformer
motion jitters
motion stability index

期刊

IEEE Journal of Selected Topics in Signal Processing 封面图
IEEE Journal of Selected Topics in Signal Processing
IF:
13.7
论文数:
1.9K
被引数:
1.1W

机构

S
shanghai jiao tong university
学者数:
15.7W
论文数: 11.7W
被引数: 159
T
tsinghua university
学者数:
11.9W
论文数: 10.0W
被引数: 137
引用论文

引用论文

Deep Online Video Stabilization Using IMU Sensors
err2023-01-01
err5
PREAI
errLi, Chen; Song, Li; Chen, Shuai; Xie, Rong; Zhang, Wenjun
err分享
err收藏
Digital urban production: how does Industry 4.0 reconfigure productive value creation in urban contexts?
err2021-08-19
err0
errOAAI
errHans-Christian Busch; Caroline Mühl; Martina Fuchs; Martina Fromhold-Eisebith
err分享
err收藏
err分享
err收藏
Text-based Editing of Talking-head Video基于文本的说话视频编辑
err2019-07-12
err171
errOAAI
errFried, Ohad; Tewari, Ayush; Zollhofer, Michael; Finkelstein, Adam; Shechtman, Eli; Goldman, Dan B.; Genova, Kyle; Jin, Zeyu; Theobalt, Christian; Agrawala, Maneesh
err分享
err收藏
Learning a model of facial shape and expression from 4D scans从4D扫描中学习面部形状和表情模型
err2017-11-20
err413
PREAI
errLi, Tianye; Bolkart, Timo; Black, Michael J.; Li, Hao; Romero, Javier
err分享
err收藏
Deep Online Video Stabilization With Multi-Grid Warping Transformation Learning基于多网格扭曲变换学习的深度在线视频稳像
err2019-05-01
err87
PREAI
errWang, Miao; Yang, Guo-Ye; Lin, Jin-Kun; Zhang, Song-Hai; Shamir, Ariel; Lu, Shao-Ping; Hu, Shi-Min
err分享
err收藏
学者 查看更多内容