arrow
Return

Explanation-Guided Adversarial Example Attacks

delete2024-05-01
delete0
PRE
AI
A
Anli Yan
X
Xiaozhang Liu *
W
Wanman Li
H
Hongwei Ye
L
Lang Li
DOI:10.1016/j.bdr.2024.100451delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Neural network -based classifiers are vulnerable to adversarial example attacks even in a black -box setting. Existing adversarial example generation technologies mainly rely on optimization -based attacks, which optimize the objective function by iterative input perturbation. While being able to craft adversarial examples, these techniques require big budgets. Latest transfer -based attacks, though being limited queries, also have a disadvantage of low attack success rate. In this paper, we propose an adversarial example attack method called MEAttack using the model -agnostic explanation technology, which can more efficiently generate adversarial examples in the black -box setting with limited queries. The core idea is to design a novel model -agnostic explanation method for target models, and generate adversarial examples based on model explanations. We experimentally demonstrate that MEAttack outperforms the state-of-the-art attack technology, i.e., AutoZOOM. The success rate of MEAttack is 4.54%-47.42% higher than AutoZOOM, and its query efficiency is reduced by 2.6-4.2 times. Experimental results show that MEAttack is efficient in terms of both attack success rate and query efficiency.
Keywords:
Deep neural network
Model explanation
Adversarial examples
Black-box
Label-only

Journal

Big Data Research cover
Big Data Research
IF:
4.2
Papers:
416
Citations:
1.1K

Organization

H
Hainan University
Scholars:
2.0W
Papers: 1.2W
Citations: 1.9W
Cited Papers

Cited Papers

errShare
errSave
Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization
err2019-10-11
err8.0K
errOAAI
errSelvaraju, Ramprasaath R.; Cogswell, Michael; Das, Abhishek; Vedantam, Ramakrishna; Parikh, Devi; Batra, Dhruv
errShare
errSave
Ensemble adversarial black-box attacks against deep learning systems
err2020-05-01
err39
PREAI
errHang, Jie; Han, Keji; Chen, Hui; Li, Yun
errShare
errSave