arrow
返回

Still-Moving: Customized Video Generation without Customized Video Data

delete2024-11-19
delete0
delete
OA
AI
H
Hila Chefer *
R
Roni Paiss
O
Omer Tov
M
Michael Rubinstein
L
Lior Wolf
T
Tali Dekel
T
Tomer Michaeli
I
Inbar Mosseri
DOI:10.1145/3687945delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Customizing text-to-image (T2I) models has seen tremendous progress recently, particularly in areas such as personalization, stylization, and conditional generation. However, expanding this progress to video generation is still in its infancy, primarily due to the lack of customized video data. In this work, we introduce Still-Moving, a novel generic framework for customizing a text-to-video (T2V) model, without requiring any customized video data. The framework applies to the prominent T2V design where the video model is built over a T2I model (e.g., via inflation). We assume access to a customized version of the T2I model, trained only on still image data (e.g., using DreamBooth). Naively plugging in the weights of the customized T2I model into the T2V model often leads to significant artifacts or insufficient adherence to the customization data. To overcome this issue, we train lightweight Spatial Adapters that adjust the features produced by the injected T2I layers. Importantly, our adapters are trained on frozen videos (i.e., repeated images), constructed from image samples generated by the customized T2I model. This training is facilitated by a novel Motion Adapter module, which allows us to train on such static videos while preserving the motion prior of the video model. At test time, we remove the Motion Adapter modules and leave in only the trained Spatial Adapters. This restores the motion prior of the T2V model while adhering to the spatial prior of the customized T2I model. We demonstrate the effectiveness of our approach on diverse tasks including personalized, stylized, and conditional generation. In all evaluated scenarios, our method seamlessly integrates the spatial prior of the customized T2I model with a motion prior supplied by the T2V model.
Keyword:
Video Customization
Video Personalization
Video Stylization
Video Editing
Video Diffusion Models

期刊

ACM Transactions on Graphics 封面图
ACM Transactions on Graphics
IF:
9.5
论文数:
4.7K
被引数:
3.6W

机构

W
Weizmann Institute of Science
学者数:
1.3W
论文数: 1.1W
被引数: 2.3W
T
Technion Israel Institute of Technology
学者数:
1.6W
论文数: 1.5W
被引数: 2.0W
T
Tel Aviv University
学者数:
3.7W
论文数: 3.0W
被引数: 3.6W
学者 查看更多机构
引用论文

引用论文

Intra-Abdominal Splenosis Mimicking Metastatic Cancer
err2011-03-01
err0
PREAI
errNicholas J. Short; Teresa G. Hayes; Peeyush Bhargava
err分享
err收藏
A Neural Space-Time Representation for Text-to-Image Personalization
err2023-12-05
err3
errOAAI
errAlaluf, Yuval; Richardson, Elad; Metzer, Gal; Cohen-Or, Daniel
err分享
err收藏
err分享
err收藏
Controlled incandescence
err2011-10-12
err0
PREAI
errJean-Jacques Greffet
err分享
err收藏
Dose rate effects on damage accumulation and void growth in self-ion irradiated tungsten
err2021-07-01
err0
errOAAI
errWeilin Jiang; Yuanyuan Zhu; Limin Zhang; Danny J. Edwards; Nicole R. Overman; Giridhar Nandipati; Wahyu Setyawan; Charles H. Henager; Richard J. Kurtz
err分享
err收藏
没有更多内容