Return
Diffusion-Based Adversarial Purification With Feature Distillation
DOI:10.1109/TBDATA.2025.3594243.png)
Abstract
En 中文
Adversarial purification is a defense strategy that utilizes generative models to neutralize adversarial perturbations. Diffusion models stand out for their powerful generative ability, making them the latest choice for generative models in adversarial purification methods. It is crucial to keep the balance between robustness against attacks and the integrity of the image content. However, these diffusion models are typically trained exclusively on clean data. We experimentally demonstrate that in the adversarial purification task, the data distribution operated by the diffusion model is often contaminated by adversarial perturbations. Adjusting the diffusion length to effectively mask adversarial perturbations while preserving image label semantics proves challenging. In this paper, we propose an innovative distillation-based diffusion approach for adversarial purification, enabling the diffusion model to operate effectively on the contaminated data distribution. Our approach involves a novel training process that integrates adversarial samples into the diffusion model's training by considering adversarial perturbations as part of the diffusion noise predicted by the model. The key to resisting perturbations is to align the feature representation of the purified image more closely with that of the clean image. To accomplish this, we develop a feature distillation technique to empower the diffusion model to learn and extract clean feature representations from adversarial samples. Extensive experimentation shows that our method achieves state-of-the-art performance against various adaptive attack benchmarks.
Keywords:
Adversarial purification
diffusion model
feature distillation
robustness
Journal
I
IF:
5.7
Papers:
860
Citations:
3.0K

