arrow
Return

DiffPortraitVideo: Diffusion-Based Expression-Consistent Zero-Shot Portrait Video Translation

delete2025-12-10
delete0
PRE
AI
S
Shaoxu Li
C
Chuhang Ma
Y
Ye Pan
DOI:10.1109/TVCG.2025.3642300delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Zero-shot text-to-video diffusion models are crafted to expand pre-trained image diffusion models to the video domain without additional training. In recent times, prevailing techniques commonly rely on existing shapes as constraints and introduce inter-frame attention to ensure texture consistency. However, such shape constraints tend to restrict the stylized geometric deformation of videos and inadvertently neglect the original texture characteristics. Furthermore, existing methods suffer from flickering and inconsistent facial expressions. In this paper, we present DiffPortraitVideo. The framework employs a diffusion model-based feature and attention injection mechanism to generate key frames, with cross-frame constraints to enforce coherence and adaptive feature fusion to ensure expression consistency. Our approach achieves high spatio-temporal and expression consistency while retaining the textual and original image properties. Extensive and comprehensive experiments are conducted to validate the efficacy of our proposed framework in generating personalized, high-quality, and coherent videos. This not only showcases the superiority of our method over existing approaches but also paves the way for further research and development in the field of text-to-video generation with enhanced personalization and quality.
Keywords:
Portrait video translation
2D diffusion models
attention mechanism

Journal

IEEE Transactions on Visualization and Computer Graphics cover
IEEE Transactions on Visualization and Computer Graphics
IF:
6.5
Papers:
337
Citations:
2.2W

Organization

S
shanghai jiao tong university
Scholars:
15.7W
Papers: 11.7W
Citations: 159
Cited Papers

Cited Papers

Control4D: Efficient 4D Portrait Editing With Text
err2024-06-16
err0
PREAI
errShao,Ruizhi; Sun,Jingxiang; Peng,Cheng; Zheng,Zerong; Zhou,Boyao; Zhang,Hongwen; Liu,Yebin
errShare
errSave
errShare
errSave
Portrait Video Editing Empowered by Multimodal Generative Priors
err2024-12-03
err0
PREAI
errGao,Xuan; Xiao,Haiyao; Zhong,Chenglai; Hu,Shimin; Guo,Yudong; Zhang,Juyong
errShare
errSave
High-Resolution Image Synthesis with Latent Diffusion Models
err2022-06-01
err0
errOAAI
errRobin Rombach; Andreas Blattmann; Dominik Lorenz; Patrick Esser; Bjorn Ommer
errShare
errSave
FateZero: Fusing Attentions for Zero-shot Text-based Video Editing
err2023-10-01
err0
errOAAI
errChenyang Qi; Xiaodong Cun; Yong Zhang; Chenyang Lei; Xintao Wang; Ying Shan; Qifeng Chen
errShare
errSave
researcher View more