Return
Interpretability adversarial example detection through multi-interpretation aggregation comparison
DOI:10.1016/j.asoc.2025.113212.png)
Abstract
En 中文
Deep neural network adversarial attacks have extended to the domain of interpretability, resulting in interpretation manipulation. Therefore, we propose a novel approach for detecting explanatory adversarial examples based on multi-interpretation aggregation comparison, called AGGEC (Aggregate Explain Compare). Our approach employs the aggregation of multiple different interpretation results to enable a comprehensive analysis of the disparities in texture shapes before and after the interpretation aggregation and employs gray level co-occurrence matrix to extract texture features before and after the interpretation aggregation to accentuate the discernible distinctions. These dissimilarities are subsequently utilized to train an external detector. AGGEC exhibits exceptional detection performance in black-box, gray-box, and white-box scenarios, achieving detection success rates of 99.4% and 95.6% on CIFAR-10 and ImageNet, respectively. Empirical results substantiate the efficacy of our proposed method in effectively identifying malicious manipulation of interpretation results.
Keywords:
Deep learning
Multi-interpretation aggregation
Adversarial example detection
Saliency map
Aggregate Explain Compare
Texture features
Journal
IF:
6.6
Papers:
1.4W
Citations:
4.8W

