Return
Improving attribution through transferable adversarial attacks
DOI:10.1016/j.patcog.2026.114496.png)
Abstract
En 中文
The interpretability of deep neural networks is crucial for understanding model decisions in various applications, including computer vision. In this paper, we propose AttEXplore+, a unified adversarial attribution framework built upon AttEXplore that connects transferable adversarial exploration with gradient-path attribution. Rather than treating transferable attacks merely as additional attack choices, AttEXplore+ reformulates them as gradient acquisition operators for constructing smoother and more structured decision-boundary exploration paths. We instantiate AttEXplore++ with validated operators such as MIG and GRA, and conduct extensive experiments on five models, including CNNs (Inception-v3, ResNet-50, VGG16) and vision transformers (MaxViT-T, ViT-B/16), using the ImageNet dataset. Our method achieves an average performance improvement of 7.57% over AttEXplore and 32.62% compared to other state-of-the-art interpretability algorithms. Using insertion and deletion scores as evaluation metrics, we show that adversarial transferability plays a vital role in enhancing attribution results. Furthermore, we explore the impact of randomness, perturbation rate, noise amplitude, and diversity probability on attribution performance, demonstrating that AttEXplore++ provides more stable and reliable explanations across various models. We release our code at: https://github.com/KxPlaug/ATTEXPLOREP.
Keywords:
Interpretability
Transferable adversarial attack
Explainable AI
Attribution
Journal
IF:
7.6
Papers:
1.3W
Citations:
4.5W
Organization
Cited Papers
No cited papers available

