arrow
Return

Armor: Shielding Unlearnable Examples Against Data Augmentation

delete2026-01-12
delete0
PRE
AI
X
Xueluan Gong
Y
Yuji Wang
Y
Yanjiao Chen
H
Haocheng Dong
Y
Yiming Li
S
Sun
S
Shuaike Li
Q
Qian Wang
DOI:10.1109/TPAMI.2026.3652456delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Private data, when published online, may be collected by unauthorized parties to train deep neural networks (DNNs). To protect privacy, defensive noises can be added to original samples to degrade their learnability by DNNs. Recently, unlearnable examples (Huang et al., 2021) are proposed to minimize the training loss such that the model learns almost nothing. However, raw data are often pre-processed before being used for training, which may restore the private information of protected data. In this paper, we reveal the data privacy violation induced by data augmentation, a commonly used data pre-processing technique to improve model generalization capability, which is the first of its kind as far as we are concerned. We demonstrate that data augmentation can significantly raise the accuracy of the model trained on unlearnable examples from 21.3% to 66.1%. To address this issue, we propose a defense framework, dubbed Armor, to protect data privacy from potential breaches of data augmentation. To overcome the difficulty of having no access to the model training process, we design a non-local module-assisted surrogate model that better captures the effect of data augmentation. In addition, we design a surrogate augmentation selection strategy that maximizes distribution alignment between augmented and non-augmented samples, to choose the optimal augmentation strategy for each class. We also use a dynamic step size adjustment algorithm to enhance the defensive noise generation process. Extensive experiments are conducted on 4 datasets and 5 data augmentation methods to verify the performance of Armor. Comparisons with 6 state-of-the-art defense methods have demonstrated that Armor can preserve the unlearnability of protected private data under data augmentation. Armor reduces the test accuracy of the model trained on augmented protected samples by as much as 60% more than baselines. We also show that Armor is robust to adversarial training. We will open-source our codes upon publication.
Keywords:
Unlearnable examples
data augmentation
and data privacy preservation

Journal

IEEE Transactions on Pattern Analysis and Machine Intelligence cover
IEEE Transactions on Pattern Analysis and Machine Intelligence
IF:
18.6
Papers:
864
Citations:
9.8W

Organization

W
Wuhan University
Scholars:
5.0K
Papers: 1.7K
Citations: 10.0W
N
nanyang technological university
Scholars:
2.5K
Papers: 1.6K
Citations: 1
W
wuhan university
Scholars:
8.1W
Papers: 5.8W
Citations: 70
Z
zhejiang university
Scholars:
17.7W
Papers: 12.1W
Citations: 152
researcher View more organizations