1
Return

Safe and Reliable Diffusion Models via Subspace Projection

delete2026-05-12
delete0
PRE
AI
H
Huiqiang Chen
T
Tianqing Zhu
L
Linlin Wang
Y
Yu Xin
L
Longxiang Gao
W
Wanlei Zhou
DOI:10.1109/tdsc.2026.3692493delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Large-scale text-to-image (T2I) diffusion models have revolutionized image generation, enabling the synthesis of highly detailed visuals from textual descriptions. However, these models may inadvertently generate inappropriate content, such as copyrighted works or offensive images. While existing methods attempt to eliminate specific unwanted concepts, they often fail to ensure robust removal—allowing the concept to reappear in subtle forms. For instance, a model may successfully avoid generating images in Van Gogh’s style when explicitly prompted with “Van Gogh”, yet still reproduce his signature artwork when given the prompt “Starry Night”. In this paper, we propose SAFER, a novel and efficient approach for thoroughly removing target concepts from diffusion models. At a high level, SAFER is inspired by the observed low-dimensional structure of the text embedding space. The method first identifies a concept-specific subspace <inline-formula><tex-math notation="LaTeX">$\mathcal{S}_{c}$</tex-math></inline-formula> associated with the target concept <inline-formula><tex-math notation="LaTeX">$c$</tex-math></inline-formula>. It then projects the prompt embeddings onto the complementary subspace of <inline-formula><tex-math notation="LaTeX">$\mathcal{S}_{c}$</tex-math></inline-formula>, effectively erasing the concept from the generated images. Since concepts can be abstract and difficult to fully capture using natural language alone, we employ textual inversion to learn an optimized embedding of the target concept from a reference image. This enables more precise subspace estimation and enhances removal performance. Furthermore, we introduce a subspace expansion strategy to ensure comprehensive and robust concept erasure. Extensive experiments demonstrate that SAFER consistently and effectively erases unwanted concepts from diffusion models while preserving generation quality.
Keywords:
Diffusion models
text-to-image generation
multimodal learning
trustworthy machine learning

Journal

IEEE Transactions on Dependable and Secure Computing cover
IEEE Transactions on Dependable and Secure Computing
IF:
7.5
Papers:
2.4K
Citations:
9.6K

Organization

A
adelaide university
Scholars:
3.5K
Papers: 1.6K
Citations: 1
C
City University of Macau
Scholars:
402
Papers: 281
Citations: 2.5K
Cited Papers

Cited Papers

Citing Papers

Citing Papers