返回
StyleCrafter: Taming Stylized Video Diffusion withReference-Augmented Adapter Learning
DOI:10.1145/3687975.png)
摘要
En 中文
Text-to-video (T2V) models have shown remarkable capabilities in gen-erating diverse videos. However, they struggle to produce user-desiredartistic videos due to (i) text's inherent clumsiness in expressing specific styles and (ii) the generally degraded style fidelity. To address these chal-lenges, we introduce StyleCrafter, a generic method that enhances pre-trained T2V models with a style control adapter, allowing video genera-tion in any style by feeding a reference image. Considering the scarcity ofartistic video data, we propose to first train a style control adapter usingstyle-rich image datasets, then transfer the learned stylization ability tovideo generation through a tailor-made finetuning paradigm. To promotecontent-style disentanglement, we employ carefully designed data augmen-tation strategies to enhance decoupled learning. Additionally, we proposea scale-adaptive fusion module to balance the influences of text-based con-tent features and image-based style features, which helps generalizationacross various text and style combinations. StyleCrafter efficiently generateshigh-quality stylized videos that align with the content of the texts andresemble the style of the reference images. Experiments demonstrate thatour approach is more flexible and efficient than existing competitors.
Keyword:
Diffusion Model
Stylized Generation
Image/Video Synthesis
期刊
IF:
9.5
论文数:
4.7K
被引数:
3.6W
机构
引用论文
Sensorless Control of Z Source Inverter fed BLDC Motor Drive by FOC - DTC Hybrid Control Strategy Using Fuzzy Logic Controller采用模糊逻辑控制器的foc-dtc混合控制策略的Z源逆变器馈电BLDC电机驱动的无传感器控制

