1
Return

SPDHead: Personalized talking head generation with semantic-guided texture compensation and pose-aware control☆

delete2026-05-23
delete0
PRE
AI
W
Wengang Zhong
Y
Yanwen Wang
W
Weimin Lei
W
Wei Zhang *
X
Xinyi Chen
DOI:10.1016/j.displa.2026.103444delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recent advancements in talking head generation have significantly improved animation quality and visual realism. However, existing methods typically require large-scale identity-specific data, making few-shot adaptation particularly challenging-especially in cross-identity scenarios where maintaining semantic consistency and preserving fine-grained facial details is essential. To address these limitations, we present SPDHead, a two-stage diffusion-based framework for personalized talking head synthesis from a small number of reference images. In the first stage, SPDHead learns general facial priors by training a conditional diffusion model on the FFHQ dataset. In the second stage, it performs fast identity adaptation via lightweight fine-tuning on as few as 20 images per subject, enabling personalized and consistent generation. To enhance texture representation beyond conventional 3D priors, we introduce a Semantic-Guided Texture Compensation (SGTC) module, which fuses global semantic embeddings from CLIP and local facial parsing features via cross-attention to enrich identity appearance modeling. Additionally, we propose a Pose-Aware Control Adapter (PACA) that projects high-dimensional FLAME coefficients into the diffusion model's latent space via a gated modulation mechanism, enabling accurate and stable control of pose and expression. Extensive experiments demonstrate that SPDHead achieves superior identity preservation, expression controllability, and detail fidelity across a wide range of motion conditions. Moreover, the framework supports intuitive portrait editing by directly manipulating semantic motion parameters at inference time, offering flexibility for interactive applications.
Keywords:
Talking head generation
Diffusion models
Few-shot learning
Semantic-guided texture compensation
Pose-aware control

Journal

Displays cover
Displays
IF:
3.4
Papers:
2.1K
Citations:
3.2K

Organization

N
northeastern university - china
Scholars:
3.0W
Papers: 2.7W
Citations: 37
Cited Papers

Cited Papers

Citing Papers

Citing Papers