Return
Provably Noise-to-Noise robust learning: Mitigating data poisoning attacks against classifier-guided diffusion models
Y
L
Z
J
H
DOI:10.1016/j.patcog.2026.113726.png)
Abstract
En 中文
Classifier-guided diffusion models generate highly realistic images by using a classifier to steer the sampling trajectory. However, recent studies indicate that diffusion pipelines are vulnerable to training-time data poisoning. In real-world deployments, users often adopt off-the-shelf pretrained models and cannot audit their training data or verify that the released checkpoints are clean. To address this threat, we propose the Noise-to-Noise method, a provable and plug-and-play safety wrapper for pretrained classifier-guided diffusion models. It provides robustness against clean-label data poisoning without retraining the base model from scratch. Specifically, the method adds a classifier-guided denoising stage before classification to purify potentially poisoned inputs. It also injects Gaussian noise to enhance robustness at inference time while maintaining accuracy on benign samples. We provide a theoretical analysis and derive a robustness bound for classifier-guided diffusion models under randomised denoised smoothing. We further implement purification sampling by applying a smoothing denoiser in the reverse diffusion process, and we evaluate the resulting denoised smoothing classifier under multiple poisoning attacks. Extensive experiments show that our method provides a stronger robustness-accuracy trade-off than Adversarial Neuron Pruning and Inference-time Clipping, while better preserving image semantics.
Keywords:
Classifier-guided diffusion models
Poisoning perturbation noise sampling
Data poisoning defence
Denoised smoothing classifier
Journal
IF:
7.6
Papers:
1.3W
Citations:
4.5W

