arrow
Return

Towards Interpretable Adversarial Examples via Sparse Adversarial Attack

delete2026-01-01
delete0
PRE
AI
F
Fudong Lin
J
Jiadong Lou
王浩 cover
王浩 (Hao Wang)
B
Brian Jalaian
X
Xu Yuan *
DOI:10.1007/978-3-032-06109-6_6delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Sparse attacks are to optimize the magnitude of adversarial perturbations for fooling deep neural networks (DNNs) involving only a few perturbed pixels (i.e., under the l(0) constraint), suitable for interpreting the vulnerability of DNNs. However, existing solutions fail to yield interpretable adversarial examples due to their poor sparsity. Worse still, they often struggle with heavy computational overhead, poor transferability, and weak attack strength. In this paper, we aim to develop a sparse attack for understanding the vulnerability of DNNs by minimizing the magnitude of initial perturbations under the l(0) constraint, to overcome the existing drawbacks while achieving a fast, transferable, and strong attack to DNNs. In particular, a novel and theoretical sound parameterization technique is introduced to approximate the NP-hard l(0) optimization problem, making directly optimizing sparse perturbations computationally feasible. Besides, a novel loss function is designed to augment initial perturbations by maximizing the adversary property and minimizing the number of perturbed pixels simultaneously. Extensive experiments are conducted to demonstrate that our approach, with theoretical performance guarantees, outperforms state-of-the-art sparse attacks in terms of computational overhead, transferability, and attack strength, expecting to serve as a benchmark for evaluating the robustness of DNNs. In addition, theoretical and empirical results validate that our approach yields sparser adversarial examples, empowering us to discover two categories of noises, i.e., obscuring noise and leading noise, which will help interpret how adversarial perturbation misleads the classifiers into incorrect predictions. Our code is available at https://github.com/fudong03/SparseAttack.
Keywords:
Sparse Attack
Adversarial Attack
Interpretability

Journal

M
MACHINE LEARNING AND KNOWLEDGE DISCOVERY IN DATABASES. RESEARCH TRACK, ECML PKDD 2025, PT VII
IF:
0
Papers:
24
Citations:
0

Organization

State University System of Florida cover
State University System of Florida
Scholars:
12.7W
Papers: 10.9W
Citations: 130
S
stevens institute of technology
Scholars:
384
Papers: 230
Citations: 0
U
university of delaware
Scholars:
1.8K
Papers: 838
Citations: 0
researcher View more organizations