Return
Probing unlearned diffusion models: A transferable adversarial attack perspective
DOI:10.1016/j.patcog.2025.112916.png)
Abstract
En 中文
Advanced text-to-image diffusion models raise safety concerns regarding identity privacy violation, copyright infringement, and Not Safe For Work (NSFW) content generation. Towards this, unlearning methods have been developed to erase these involved concepts from diffusion models. However, these unlearning methods only shift the text-to-image mapping and preserve the visual content within the generative space, leaving a fatal flaw for restoring these erased concepts. This inherent limitation necessitates probing the reliability of unlearning methods against adversarial inputs, but previous probing approaches are sub-optimal from two perspectives: (1) Lack of transferability: Some methods require white-box access to the unlearned model, and learned adversarial inputs often fail to transfer to other unlearned models for concept restoration; (2) Limited attack: Prompt-level methods struggle to restore narrow concepts like celebrity identities from unlearned models. Therefore, this paper aims to leverage the transferability of the adversarial attack to probe unlearning robustness under a gray-box setting. This challenging scenario assumes that the unlearning method is unknown and the unlearned model is inaccessible for optimization. To address this challenge, we first analyze the reasons for the poor transferability of previous methods. Then, we employ an Adversarial Search (AS) strategy to search for the adversarial embedding which can transfer across different unlearned models. This strategy adopts the original Stable Diffusion model as a surrogate model to iteratively erase and search for embeddings, enabling it to find the embedding that can restore the target concept for different unlearning methods. Extensive experiments demonstrate the transferability of the searched adversarial embedding across several state-of-the-art unlearning methods and its effectiveness for different levels of concepts, including fine-grained celebrity identities. Additionally, we explore mapping the identified embedding back to prompts, achieving comparable performance with prompt-level baselines, and demonstrate the potential for enhancing unlearning robustness using the obtained embedding.
Keywords:
Diffusion model
Machine unlearning
Adversarial attack

