arrow
返回

Speech driven video editing via an audio-conditioned diffusion model

delete2024-02-01
delete4
delete
OA
AI
D
Dan Bigioi *
B
Basak, Shubhajit
M
Michał Stypułkowski
M
Maciej Zięba
H
Hugh Jordan
R
Rachel McDonnell
P
Peter Corcoran
DOI:10.1016/j.imavis.2024.104911delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Taking inspiration from recent developments in visual generative tasks using diffusion models, we propose a method for end-to-end speech-driven video editing using a denoising diffusion model. Given a video of a talking person, and a separate auditory speech recording, the lip and jaw motions are re-synchronised without relying on intermediate structural representations such as facial landmarks or a 3D face model. We show this is possible by conditioning a denoising diffusion model on audio mel spectral features to generate synchronised facial motion. Proof of concept results are demonstrated on both single -speaker and multi-speaker video editing, providing a baseline model on the CREMA-D audiovisual data set. To the best of our knowledge, this is the first work to demonstrate and validate the feasibility of applying end-to-end denoising diffusion models to the task of audiodriven video editing. All code, datasets, and models used as part of this work are made publicly available here: https://danbigioi.github.io/DiffusionVideoEditing/.
Keyword:
Video editing
Talking head generation
Generative AI
Diffusion models
Dubbing
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Image and Vision Computing 封面图
Image and Vision Computing
IF:
4.2
论文数:
4.1K
被引数:
6.7K

机构

U
University of Wroclaw
学者数:
4.3K
论文数: 4.4K
被引数: 4.1K
O
ollscoil na gaillimhe-university of galway
学者数:
1.1W
论文数: 8.7K
被引数: 5
W
wroclaw university of science & technology
学者数:
7.4K
论文数: 7.1K
被引数: 2
学者 查看更多机构
引用论文

引用论文

You Said That?: Synthesising Talking Faces from Audio
err2019-02-13
err101
errOAAI
errJamaludin, Amir; Chung, Joon Son; Zisserman, Andrew
err分享
err收藏
err分享
err收藏
Synthesizing Obama: Learning Lip Sync from Audio
err2017-07-20
err665
PREAI
errSuwajanakorn, Supasorn; Seitz, Steven M.; Kemelmacher-Shlizerman, Ira
err分享
err收藏
Everybody's Talkin': Let Me Talk as You Want
err2022-01-01
err48
errOAAI
errSong, Linsen; Wu, Wayne; Qian, Chen; He, Ran; Loy, Chen Change
err分享
err收藏
Design of High-Performance Microprocessor Circuits
err
IF0
err2000-01-01
err0
PREAI
errAnantha Chandrakasan; William J. Bowhill; Frank Fox
err分享
err收藏
Audio-Driven Facial Animation by Joint End-to-End Learning of Pose and Emotion
err2017-07-20
err256
PREAI
errKarras, Tero; Aila, Timo; Laine, Samuli; Herva, Antti; Lehtinen, Jaakko
err分享
err收藏
Diffusion Models: A Comprehensive Survey of Methods and Applications扩散模型: 方法和应用的综合综述
err2023-11-09
err245
errOAAI
errYang, Ling; Zhang, Zhilong; Song, Yang; Hong, Shenda; Xu, Runsheng; Zhao, Yue; Zhang, Wentao; Cui, Bin; Yang, Ming-Hsuan
err分享
err收藏
Generative Adversarial Networks生成对抗网络
err2020-10-22
err1.0W
errOAAI
errGoodfellow, Ian; Pouget-Abadie, Jean; Mirza, Mehdi; Xu, Bing; Warde-Farley, David; Ozair, Sherjil; Courville, Aaron; Bengio, Yoshua
err分享
err收藏
Diffsound: Discrete Diffusion Model for Text-to-Sound Generation
err2023-01-01
err44
errOAAI
errYang, Dongchao; Yu, Jianwei; Wang, Helin; Wang, Wen; Weng, Chao; Zou, Yuexian; Yu, Dong
err分享
err收藏
学者 查看更多内容