arrow
返回

Efficient Method for Robust Backdoor Detection and Removal in Feature Space Using Clean Data

delete2025-01-01
delete0
delete
OA
AI
D
Donik Vršnak *
M
Marko Subašić
S
Sven Lončarić
DOI:10.1109/ACCESS.2025.3531716delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The steady increase of proposed backdoor attacks on deep neural networks highlights the need for robust defense methods for their detection and removal. A backdoor attack is a type of attack where hidden triggers are added to the input data during training, with the goal of changing the behavior of the model during inference. These attacks pose a significant security threat in critical applications, such as street sign or pedestrian recognition for autonomous vehicles, biometric authentication, image retrieval, semantic labeling, etc. To combat these threats, many defense mechanisms have been proposed. These methods target different areas, such as computer vision (CV), natural language processing (NLP), and thus utilize different assumptions about the nature of the input data and the type of backdoor trigger used in the attack. However, the attacker can exploit these assumptions, which reduces their successfulness in real-world scenarios. Thus, a robust method for backdoor detection needs to have broad and simple assumptions. Furthermore, detection methods that rely on the input data suffer from the fact that they are constrained to the modality of the input and cannot apply to different modalities. In this work, a novel method for backdoor detection and removal for classification tasks using features extracted by the attacked model called FEAT-IN is proposed. This method can detect and reconstruct the feature representation of the possible triggers used in attacking the neural network. Using these reconstructed trigger features, the method can be used to efficiently mitigate the effects of an attack. Extensive experiments on multiple datasets and attack methods demonstrate that, when compared to state-of-the-art methods such as Neural Cleanse, Neural Attention Distillation, I-BAU, BTI-DBF etc. the FEAT-IN method provides several benefits. It can more consistently detect and mitigate backdoor attacks than similar trigger inversion defense methods that conduct the defense in the input space instead of feature space (where, on average, it achieves approx. 10% higher decrease in attack success rate during mitigation compared to the second-best method). Secondly, it reduces the memory footprint and the computation time by at least an order of magnitude compared to other methods, which allows FEAT-IN to be used practically in real-world scenarios. Finally, it is not constrained to only computer vision tasks, as this assumption holds for feature spaces of different problems, which is demonstrated by applying it without any change to semantic analysis on the SST-2 dataset.
Keyword:
Feature extraction
Training
Prevention and mitigation
Optimization
Data models
Image reconstruction
Computer vision
Computational modeling
Visualization
Vectors
Backdoor attack
backdoor defense
image classification
neural cleanse

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

U
University of Zagreb
学者数:
1.8W
论文数: 1.3W
被引数: 1.1W
引用论文

引用论文

err分享
err收藏
Monitoring of Water Transportation in Plant Stem With Microneedle Sap Flow Sensor
err2018-06-01
err0
PREAI
errSangwoong Baek; Eunyong Jeon; Kyoung Sub Park; Kyung-Hwan Yeo; Junghoon Lee
err分享
err收藏
err
IF0
err
err0
errOAAI
err
err分享
err收藏
An Automated UAV Mission System
err
IF0
err2003-09-01
err0
PREAI
errKatheerine D. Mullens; Estrellina B. Pacis; Stephen B. Stancliff; Aaron B. Burmeister; Thomas A. Denewiler
err分享
err收藏
Utilizing non-stoichiometry in Nd2Zr2O7pyrochlore: exploring superior ionic conductors
err2016-01-01
err0
PREAI
errP. Anithakumari; V. Grover; C. Nandi; K. Bhattacharyya; A. K. Tyagi
err分享
err收藏
学者 查看更多内容