Return
Rethinking Reverse-Engineering-Based Backdoor Removal in Self-Supervised Learning
DOI:10.1109/tifs.2026.3723195.png)
Abstract
En 中文
Self-supervised learning (SSL) has garnered increasing attention due to its ability to train high-performing models without requiring labeled data, achieving remarkable results on various downstream tasks, such as traffic sign recognition. However, existing works show that self-supervised learning is vulnerable to backdoor attacks, causing potential security threats. Reverse engineering can reconstruct the implanted backdoor trigger and then remove backdoors by mitigating the influence of inverted triggers. Nevertheless, existing reverse-engineering methods either focus on supervised learning or can only invert low-quality triggers that are sufficient for detection but inadequate for completely removing backdoors. In this paper, we propose a comprehensive theoretical analysis of existing reverse-engineering methods and observe the phenomenon of trigger growth. We reveal the principle behind it, based on which we propose a new backdoor removal method specially designed for self-supervised learning, called RECEIVE. RECEIVE introduces a principled early-stop strategy based on Wasserstein distance to identify optimal inverted triggers during iterative trigger inversion, achieving effective backdoor unlearning while preserving model utility. The experiments on two state-of-the-art backdoor attacks in self-supervised learning demonstrate that RECEIVE can effectively reduce the attack success rate using only 500 images (e.g., as low as 0.5%) of the pre-training data, surpassing five baseline backdoor removal techniques.
Keywords:
Backdoor defense
pre-trained encoders
mutual information
distillation
Journal
IF:
8
Papers:
5.3K
Citations:
2.3W
Organization
Cited Papers
No cited papers available

