arrow
Return

Audio-Driven Talking Face Video Generation With Dynamic Convolution Kernels

delete2023-01-01
delete19
delete
OA
AI
Z
Zipeng Ye
M
Mengfei Xia
R
Ran Yi *
张
张举勇 (Juyong Zhang)
Yu-Kun Lai cover
Yu-Kun Lai (Yu‐Kun Lai)
X
Xuwei Huang
G
Guoxin Zhang
Y
Yong‐Jin Liu *
DOI:10.1109/TMM.2022.3142387delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In this paper, we present a dynamic convolution kernel (DCK) strategy for convolutional neural networks. Using a fully convolutional network with the proposed DCKs, high-quality talking-face video can be generated from multi-modal sources (i.e., unmatched audio and video) in real time, and our trained model is robust to different identities, head postures, and input audios. Our proposed DCKs are specially designed for audio-driven talking face video generation, leading to a simple yet effective end-to-end system. We also provide a theoretical analysis to interpret why DCKs work. Experimental results show that our method can generate high-quality talking-face video with background at 60 fps. Comparison and evaluation between our method and the state-of-the-art methods demonstrate the superiority of our method.
Keywords:
Dynamic kernel
convolutional neural network
multi-modal generation task
audio-driven talking-face generation

Journal

IEEE Transactions on Multimedia cover
IEEE Transactions on Multimedia
IF:
9.7
Papers:
4.5K
Citations:
2.4W

Organization

T
tsinghua university
Scholars:
11.9W
Papers: 10.0W
Citations: 137
S
shanghai jiao tong university
Scholars:
15.7W
Papers: 11.7W
Citations: 159
U
university of science & technology of china, cas
Scholars:
3.2W
Papers: 2.7W
Citations: 74
C
Cardiff University
Scholars:
2.7W
Papers: 2.5W
Citations: 3.5W
C
chinese academy of sciences
Scholars:
56.7W
Papers: 45.0W
Citations: 704
researcher View more organizations
Cited Papers

Cited Papers

Planar monomode optical couplers based on multimode interference effects
err1992-12-01
err0
errOAAI
errL.B. Soldano; F.B. Veerman; M.K. Smit; B.H. Verbeek; A.H. Dubost; E.C.M. Pennings
errShare
errSave
Text-based Editing of Talking-head Video
err2019-07-12
err171
errOAAI
errFried, Ohad; Tewari, Ayush; Zollhofer, Michael; Finkelstein, Adam; Shechtman, Eli; Goldman, Dan B.; Genova, Kyle; Jin, Zeyu; Theobalt, Christian; Agrawala, Maneesh
errShare
errSave
errShare
errSave
Effect of sildenafil on ocular hemodynamics in 3 months regular use
err2005-11-17
err0
errOAAI
errS O Dündar; Y Dayanir; A Topaloğlu; M Dündar; İ Koçak
errShare
errSave
Viability of ram spermatozoa in relation to the abstinence period and successive ejaculations
err2008-06-28
err0
errOAAI
errM. OLLERO; T. MUIÑO‐BLANCO; M. J. LÓPEZ‐PÉREZ; J. A. CEBRIÁN‐PÉREZ
errShare
errSave
Deep Video Portraits
err2018-07-30
err279
errOAAI
errKim, Hyeongwoo; Garrido, Pablo; Tewari, Ayush; Xu, Weipeng; Thies, Justus; Niessner, Matthias; Perez, Patrick; Richardt, Christian; Zollhofer, Michael; Theobalt, Christian
errShare
errSave
Bringing Portraits to Life
err2017-11-20
err124
PREAI
errAverbuch-Elor, Hadar; Cohen-Or, Daniel; Kopf, Johannes; Cohen, Michael F.
errShare
errSave
researcher View more