返回
Learning Reliable Visual Saliency For Model Explanations
DOI:10.1109/TMM.2019.2949872.png)
摘要
En 中文
By highlighting important features that contribute to model prediction, visual saliency is used as a natural form to interpret the working mechanism of deep neural networks. Numerous methods have been proposed to achieve better saliency results. However, we find that previous visual saliency methods are not reliable enough to provide meaningful interpretation through a simple sanity check: saliency methods are required to explain the output of non-maximum prediction classes, which are usually not ground-truth classes. For example, let the methods interpret an image of dog given a wrong class label fish as the query. This procedure can test whether these methods reliably interpret model's predictions based on existing features that appear in the data. Our experiments show that previous methods failed to pass the test by generating similar saliency maps or scattered patterns. This false saliency response can be dangerous in certain scenarios, such as medical diagnosis. We find that these failure cases are mainly due to the attribution vanishing and adversarial noise within these methods. In order to learn reliable visual saliency, we propose a simple method that requires the output of the model to be close to the original output while learning an explanatory saliency mask. To enhance the smoothness of the optimized saliency masks, we then propose a simple Hierarchical Attribution Fusion (HAF) technique. In order to fully evaluate the reliability of visual saliency methods, we propose a new task Disturbed Weakly Supervised Object Localization (D-WSOL) to measure whether these methods can correctly attribute the model's output to existing features. Experiments show that previous methods fail to meet this standard, and our approach helps to improve the reliability by suppressing false saliency responses. After observing a significant layout difference in saliency masks between real and adversarial samples. we propose to train a simple CNN on these learned hierarchical attribution masks to distinguish adversarial samples. Experiments show that our method can improve detection performance over other approaches significantly.
Keyword:
Visualization
Reliability
Predictive models
Task analysis
Perturbation methods
Backpropagation
Real-time systems
Model interpretability
adversarial example defense
visual salience
deep learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
9.7
论文数:
4.5K
被引数:
2.4W
机构
引用论文
Cell-free Protein Synthesis in an Autoinduction System for NMR Studies of Protein–Protein Interactions用于NMR研究蛋白质-蛋白质相互作用的自动诱导系统中的无细胞蛋白质合成
Do Perceptions of Competence Mediate The Relationship Between Fundamental Motor Skill Proficiency and Physical Activity Levels of Children in Kindergarten?能力的感知是否可以介导幼儿园儿童的基本运动技能熟练程度与身体活动水平之间的关系?
Nanocomposite of hexagonal β-Ni(OH)2/multiwalled carbon nanotubes as high performance electrode for hybrid supercapacitors六方 β-ni (OH)2/多壁碳纳米管复合材料作为混合超级电容器的高性能电极
Ferrocene studies III. Platinum and palladium complexes of [(dimethylamino)methyl]ferrocene二茂铁研究III.[(二甲基氨基) 甲基] 二茂铁的铂和钯配合物
Structural study of lanthanides(III) in aqueous nitrate and chloride solutions by EXAFS通过EXAFS对硝酸盐和氯化物水溶液中镧系元素 (III) 的结构研究

