arrow
返回

Adversarial example detection based on saliency map features

delete2021-09-06
delete15
PRE
AI
王
王莘 (Shen Wang) *
Y
Yuxin Gong
DOI:10.1007/s10489-021-02759-8delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In recent years, machine learning has greatly improved image recognition capability. However, studies have shown that neural network models are vulnerable to adversarial examples that make models output wrong answers with high confidence. To understand the vulnerabilities of models, we use interpretability methods to reveal the internal decision-making behaviors of models. Interpretation results reflect that the evolutionary process of nonnormalized saliency maps between clean and adversarial examples are increasingly differentiated along model hidden layers. By taking advantage of this phenomenon, we propose an adversarial example detection method based on multilayer saliency features, which can comprehensively capture the abnormal characteristics of adversarial example interpretations. Experimental results show that the proposed method can effectively detect adversarial examples based on gradient, optimization and black-box attacks, and it is comparable with the state-of-the-art methods.
Keyword:
Machine learning
Adversarial example detection
Interpretability
Saliency map

期刊

Applied Intelligence 封面图
Applied Intelligence
IF:
3.5
论文数:
7.6K
被引数:
1.7W

机构

H
harbin institute of technology
学者数:
8.0W
论文数: 6.6W
被引数: 66
引用论文

引用论文

err分享
err收藏
err2001-01-01
err0
PREAI
errN. A. Sanina; O. A. Rakova; S. M. Aldoshin; I. I. Chuev; E. G. Atovmyan; N. S. Ovanesyan
err分享
err收藏
Prediction of brittle-to-ductile transitions in polystyrene
err2003-01-01
err0
PREAI
errH.G.H. van Melick; L.E. Govaert; H.E.H. Meijer
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容