arrow
返回

Optimized Adversarial Example With Classification Score Pattern Vulnerability Removed

delete2022-01-01
delete4
delete
OA
AI
H
Hyun Kwon
K
Kyoungmin Ko
S
Sunghwan Kim *
DOI:10.1109/ACCESS.2021.3110473delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Neural networks provide excellent service on recognition tasks such as image recognition and speech recognition as well as for pattern analysis and other tasks in fields related to artificial intelligence. However, neural networks are vulnerable to adversarial examples. An adversarial example is a sample that is designed to be misclassified by a target model, although it poses no problem for recognition by humans, that is created by applying a minimal perturbation to a legitimate sample. Because the perturbation applied to the legitimate sample to create an adversarial example is optimized, the classification score for the target class has the characteristic of being similar to that for the legitimate class. This regularity occurs because minimal perturbations are applied only until the classification score for the target class is slightly higher than that for the legitimate class. Given the existence of this regularity in the classification scores, it is easy to detect an optimized adversarial example by looking for this pattern. However, the existing methods for generating optimized adversarial examples do not consider their weakness of allowing detectability by recognizing the pattern in the classification scores. To address this weakness, we propose an optimized adversarial example generation method in which the weakness due to the classification score pattern is removed. In the proposed method, a minimal perturbation is applied to a legitimate sample such that the classification score for the legitimate class is less than that for some of the other classes, and an optimized adversarial example is created with the pattern vulnerability removed. The results show that using 500 iterations, the proposed method can generate an optimized adversarial example that has a 100% attack success rate, with distortions of 2.81 and 2.23 for MNIST and Fashion-MNIST, respectively.
Keyword:
Perturbation methods
Distortion
Neural networks
Image recognition
Task analysis
Speech recognition
Licenses
Neural network
evasion attack
classification score
optimization

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

K
Konkuk University
学者数:
1.2W
论文数: 1.1W
被引数: 1.2W
引用论文

引用论文

err分享
err收藏
err分享
err收藏
学者 查看更多内容