1
Return

On Neural Network Approximation of Ideal Adversarial Attack and Convergence of Adversarial Training\ast

delete2025-12-31
delete0
PRE
AI
H
Haldar, Rajdeep *
S
Song, Qifan
DOI:10.1137/23M1590512delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Adversarial attacks are usually expressed in terms of a gradient-based operation on the input data and model; this results in heavy computations every time an attack is generated. Recent empirical works exhibit attacks that can be approximated by neural networks. This work provides a general theoretical framework to represent adversarial attacks as a trainable function without further gradient computation. We first motivate that the theoretical best attacks, under proper conditions, can be represented as smooth piecewise functions (piecewise Ho\lder functions). Then we obtain an approximation result of such functions by a neural network. Subsequently, we emulate the ideal attack process by a neural network and reduce the adversarial training to a mathematical game between an attack network and a training model (a defense network). We also obtain convergence rates of adversarial loss in terms of the sample size n for adversarial training in such a setting.
Keywords:
adversarial robustness
neural networks
machine learning
semiwhite box attacks
projected gra-dient flow
convergence rates

Journal

S
SIAM JOURNAL ON MATHEMATICS OF DATA SCIENCE
IF:
2.6
Papers:
17
Citations:
0

Organization

Purdue University System cover
Purdue University System
Scholars:
3.9W
Papers: 3.6W
Citations: 66
Cited Papers

Cited Papers

Citing Papers

Citing Papers