arrow
Return

Attention-based backdoor attacks against natural language processing models

delete2025-02-01
delete0
PRE
AI
Y
Yunchun Zhang
Q
Qi Wang
S
Shaohui Min
R
Ruifeng Zuo
F
Feiyang Huang
H
Hao Liu
S
Shaowen Yao *
DOI:10.1016/j.asoc.2025.112907delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Backdoor attacks against natural language processing (NLP) models are surging with enhanced attack success rates. However, these backdoor attacks are limited in sentence fluency, grammar errors, and stealthiness. To address these issues, this study proposes an attention-based backdoor attack that generates high-quality backdoor-poisoned samples. The proposed backdoor employs class activation mapping (CAM) to generate backdoor texts with a baseline convolutional neural network in two steps: trigger generation and trigger insertion. The trigger generation leverages high-frequency words as candidate trigger patterns that are subsequently used to generate poisoned texts with high stealthiness and effectiveness. These candidate words are then inserted into computed positions of clean texts, under a low poisoning rate, based on available positions computed with the CAM-based attention method. Through extensive experiments on five benchmark datasets, the proposed CAM-based backdoor attack demonstrates a more excellent performance than the other five backdoor attacks from multiple aspects, including utility, effectiveness, and stealthiness. The proposed method is more robust than other attacks because it maintains almost the highest attack success rate against four benchmark defense methods.
Keywords:
Backdoor attack
Natural language processing
Class activation mapping
Trigger patterns
Attack success rates

Journal

Applied Soft Computing cover
Applied Soft Computing
IF:
6.6
Papers:
1.4W
Citations:
4.8W

Organization

E
engn res ctr cyberspace
Scholars:
4
Papers: 2
Citations: 0