Return
Sketch-text-driven rectified flow for identity-preserving 4D face generation
B
刘
W
G
J
Z
DOI:10.1007/s00371-026-04654-0.png)
Abstract
En 中文
The limitations of single-modality control in preserving facial identity and describing temporal expressions have motivated sketch-text dual-driven 4D face generation, which provides a flexible paradigm for dynamic digital face synthesis in applications requiring precise identity customization and controllable expression manipulation. However, this task remains challenging due to the synthetic-to-real domain gap in sparse sketches, cross-modal interference between heterogeneous conditions, and the scarcity of paired sketch-text 4D mesh data. To address these challenges, we propose sketch-text-driven rectified flow (STDRF), a conditional rectified-flow framework for identity-preserving and semantically controllable 4D face generation. STDRF learns a velocity field that transports Gaussian noise to the target 4D facial motion manifold under sketch-structural and text-semantic dual guidance. To obtain robust identity priors from sparse sketches, we design a sketch encoder enhanced by Geometric Contour and Texture Detail (GCTD) preprocessing and MixStyle domain adaptation. To reduce cross-modal interference, a dual-path independent cross-attention module based on IP-Adapter is designed to inject sketch features and text semantics into an Attention DiffusionNet Block (ADNB)-based denoising backbone in parallel. Furthermore, a multimodal triplet dataset is constructed by pairing 4D facial mesh sequences with synthetic sketches and hierarchical text descriptions. On the subject-independent test split, STDRF achieves an MVE of 0.81 $$\times $$ $$10^{-3}$$ , an ID-Sim of 0.989, and a GPU inference speed of 21.89 FPS for 160-frame sequences, showing favorable geometric fidelity, identity consistency, and efficiency compared with representative cascaded methods. Code is available at: https://github.com/alex11782/STDRF .
Keywords:
4D face animation
Rectified flow
Sketch-guided generation
Text-conditioned generation
Digital humans
Mesh animation
Journal
IF:
2.9
Papers:
4.5K
Citations:
6.5K
