Return
On Neural Network Approximation of Ideal Adversarial Attack and Convergence of Adversarial Training\ast
H
S
DOI:10.1137/23M1590512.png)
Abstract
En 中文
Adversarial attacks are usually expressed in terms of a gradient-based operation on the input data and model; this results in heavy computations every time an attack is generated. Recent empirical works exhibit attacks that can be approximated by neural networks. This work provides a general theoretical framework to represent adversarial attacks as a trainable function without further gradient computation. We first motivate that the theoretical best attacks, under proper conditions, can be represented as smooth piecewise functions (piecewise Ho\lder functions). Then we obtain an approximation result of such functions by a neural network. Subsequently, we emulate the ideal attack process by a neural network and reduce the adversarial training to a mathematical game between an attack network and a training model (a defense network). We also obtain convergence rates of adversarial loss in terms of the sample size n for adversarial training in such a setting.
Keywords:
adversarial robustness
neural networks
machine learning
semiwhite box attacks
projected gra-dient flow
convergence rates
Journal
S
IF:
2.6
Papers:
17
Citations:
0
