返回
Adversarial example detection based on saliency map features
DOI:10.1007/s10489-021-02759-8.png)
摘要
En 中文
In recent years, machine learning has greatly improved image recognition capability. However, studies have shown that neural network models are vulnerable to adversarial examples that make models output wrong answers with high confidence. To understand the vulnerabilities of models, we use interpretability methods to reveal the internal decision-making behaviors of models. Interpretation results reflect that the evolutionary process of nonnormalized saliency maps between clean and adversarial examples are increasingly differentiated along model hidden layers. By taking advantage of this phenomenon, we propose an adversarial example detection method based on multilayer saliency features, which can comprehensively capture the abnormal characteristics of adversarial example interpretations. Experimental results show that the proposed method can effectively detect adversarial examples based on gradient, optimization and black-box attacks, and it is comparable with the state-of-the-art methods.
Keyword:
Machine learning
Adversarial example detection
Interpretability
Saliency map
期刊
IF:
3.5
论文数:
7.6K
被引数:
1.7W
机构
引用论文
Welding characteristics of aluminum, copper, nickel and aluminum alloy with alumina coating using ultrasonic complex vibration welding equipments铝、铜、镍及铝合金氧化铝涂层超声复合振动焊接特性研究
Cell-free Protein Synthesis in an Autoinduction System for NMR Studies of Protein–Protein Interactions用于NMR研究蛋白质-蛋白质相互作用的自动诱导系统中的无细胞蛋白质合成
Do Perceptions of Competence Mediate The Relationship Between Fundamental Motor Skill Proficiency and Physical Activity Levels of Children in Kindergarten?能力的感知是否可以介导幼儿园儿童的基本运动技能熟练程度与身体活动水平之间的关系?
Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey计算机视觉中对抗性攻击对深度学习的威胁: 一项调查
IEEE ACCESS
IF3.6
Oxidative stress in liver of grass carp Ctenopharyngodon idella naturally infected with Saprolegnia parasitica and its influence on disease pathogenesis天然感染水蚤的草鱼肝脏氧化应激及其对疾病发病机制的影响

