返回
AdvCheck: Characterizing adversarial examples via local gradient checking
DOI:10.1016/j.cose.2023.103540.png)
摘要
En 中文
Deep neural networks (DNNs) are vulnerable to adversarial examples, which may lead to catastrophe in security-critical domains. Numerous detection methods are proposed to characterize the feature uniqueness of adversarial examples, or to distinguish DNN's behavior activated by the adversarial examples. Detections based on features may be compromised when faced with stronger attacks. Besides, they require a large amount of specific adversarial examples. Another mainstream, model-based detections, which characterize input properties by model behaviors, suffer from heavy computation cost. To address the issues, we introduce the concept of local gradient, and reveal that adversarial examples have a quite larger bound of local gradient than the benign ones. Inspired by the observation, we leverage local gradient for detecting adversarial examples, and propose a general framework AdvCheck. Specifically, by calculating the local gradient from a few benign examples and noise-added misclassified examples to train a detector, adversarial examples and even misclassified natural inputs can be precisely distinguished from benign ones. Through extensive experiments, we have validated the AdvCheck's superior performance to the state-of-the-art (SOTA) baselines, with detection rate (similar to x1.2) on general adversarial attacks and (similar to x1.4) on misclassified natural inputs on average, with average 1/200 time cost. We also provide interpretable results for successful detection.
Keyword:
Adversarial attack
Adversarial detection
Local gradient
Deep neural network
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
C
IF:
5.4
论文数:
4.6K
被引数:
1.4W
机构
引用论文
Misleading attention and classification: An adversarial attack to fool object detection models in the real world误导性关注和分类: 在现实世界中欺骗对象检测模型的对抗性攻击
COMPUTERS & SECURITY
IF5.4
Detecting Adversarial Image Examples in Deep Neural Networks with Adaptive Noise Reduction基于自适应降噪的深度神经网络对抗图像样本检测
FineFool: A novel DNN object contour attack on image recognition based on the attention perturbation adversarial technique
COMPUTERS & SECURITY
IF5.4
An efficient network behavior anomaly detection using a hybrid DBN-LSTM network使用混合dbn-lstm网络的高效网络行为异常检测
COMPUTERS & SECURITY
IF5.4

