arrow
Return

Prompt-Guided Semantic Latent Direction Learning in Diffusion Models for Abstract Visual Concept Manipulation

delete2026-06-25
delete0
delete
OA
AI
M
Mahzaib Khalid *
F
Fangli Ying
A
Al-Garadi Ahmed Mohammed Atef
A
Aniwat Phaphuangwittayakul
R
Riyad Dhuny
DOI:10.3390/jimaging12070279delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Diffusion-based generative models achieve high-fidelity image synthesis; however, controlling internal representations for abstract visual concepts remains challenging due to the ambiguity of textual descriptions. In this work, we propose a prompt-guided concept-vector learning framework for the controllable manipulation of such concepts without requiring external human-annotated image pairs, segmentation masks, identity labels, or manually annotated editing targets. The method introduces a learnable concept vector optimized in the bottleneck (mid-block) feature space of a pretrained Stable Diffusion U-Net, while keeping all pretrained model parameters frozen. A multi-prompt data generation strategy based on paired positive and neutral prompts provides weak semantic guidance for capturing the target concept direction and reducing dependence on a single prompt formulation. The learned vector is further applied in an image-to-image setting through controlled noise injection and concept-guided denoising, enabling the semantic modification of real images while preserving structural content. The concept strength is controlled by a scaling parameter γ , while the image-to-image noise strength is controlled by β , allowing for a practical balance between semantic modification and structural fidelity. Experiments are conducted on two main abstract concepts, perfect skin and peaceful lake, with additional qualitative analysis on subjective portrait-level concepts. Quantitative evaluation using SSIM, LPIPS, and CLIP similarity demonstrates that the proposed method improves semantic alignment while maintaining structural preservation compared with Stable Diffusion image-to-image baselines. A human preference study further shows that concept-injected outputs are preferred in 76.0% of responses for perfect skin and 85.7% for peaceful lake. Ablation studies further demonstrate the controllability and robustness of the proposed framework. Overall, the method provides a simple and parameter-efficient approach for interpretable concept-level manipulation in diffusion models.
Keywords:
diffusion models
stable diffusion
concept-vector learning
prompt-guided learning
semantic manipulation
image-to-image editing
bottleneck feature injection
abstract visual concepts

Journal

J
Journal of Imaging
IF:
3.3
Papers:
978
Citations:
4.4K

Organization

U
university of technology
Scholars:
171
Papers: 91
Citations: 0
C
Chiang Mai University
Scholars:
1.5W
Papers: 9.2K
Citations: 7.9K
E
east china university of science and technology
Scholars:
7.9K
Papers: 2.6K
Citations: 3
researcher View more organizations