arrow
Return

DMV-CLIP: Disentangled Multimodal Visual Adaptation for Text-Driven Face Editing

delete2026-06-29
delete0
PRE
AI
M
Minghao Li
F
Fan Zhang
X
Xin Wei
H
Huan Wan
H
Haoruo Zhang
X
Xuhui Huang *
DOI:10.1111/exsy.70345delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Text-driven face editing has attracted widespread interest due to its intuitive control and user-friendly interaction. However, current state-of-the-art (SOTA) methods face two main challenges: (1) they utilize unfinetuned general image-text encoders for modality fusion, making it difficult to comprehend domain-specific knowledge in facial attribute editing (dozens of fine-grained facial attributes such as moustache and lipsticks); (2) they roughly optimize all attributes simultaneously using a cross-entropy loss, leading to severe mutual interference among attributes. To this end, we propose Disentangled Multimodal Visual Adaptation for CLIP (DMV-CLIP). First, DMV-CLIP incorporates learnable context tokens to inject facial domain knowledge into the CLIP model via multimodal prompt learning (MPL). Second, it employs directional contrastive learning (DCL) to disentangle facial attributes and enable precise editing. Finally, DMV-CLIP utilizes a vision-language consistency model (VLCM) to maintain identity consistency while ensuring that the generated images strictly adhere to the semantic instructions.
Keywords:
contrastive learning
face editing
generative adversarial network
image manipulation
multimodal prompt learning

Journal

Expert Systems cover
Expert Systems
IF:
2.3
Papers:
2.5K
Citations:
3.8K

Organization

N
nanchang university
Scholars:
8.5K
Papers: 2.3K
Citations: 0
J
jiangxi normal university
Scholars:
1.2K
Papers: 413
Citations: 0
N
Nanyang Technological University
Scholars:
4.9W
Papers: 4.8W
Citations: 8.1W
researcher View more organizations