Return
Adversarial attack method based on reconstructed data and surrogate model
DOI:10.1117/1.JEI.35.1.013025.png)
Abstract
En 中文
Adversarial attacks aim to make neural network models produce incorrect outputs by adding perturbations to images. Attack methods based on surrogate models usually require access to the training data of the target model; however, obtaining this training data is often difficult due to data privacy and transmission issues. In addition, existing methods typically focus on the effectiveness of model interactions, without fully considering the performance of the generative model itself. To address these issues, we propose an adversarial attack method based on reconstructed data and a surrogate model, focusing on leveraging the output labels or probabilistic information of the target model to construct an efficient generator, which is trained under the constraints of a proposed composite loss function, without explicitly obtaining the model's training data. The reconstructed data can approximate the label distribution of the target model's training data and can then be used to train the surrogate model under another designed composite loss function constraint. Ultimately, this process enables the surrogate model to distill knowledge from the target model. Experimental results indicate that the proposed scheme can effectively perceive the target model and achieve desirable attack performance under a query budget of 250K. In particular, comparative analyses on the SVHN, CIFAR-10, and CIFAR-100 datasets against other mainstream methods demonstrate the effectiveness of this approach.
Keywords:
adversarial attack
generative model
reconstructed data
surrogate model
target model
Journal
J
IF:
1
Papers:
148
Citations:
2.7K

