arrow
Return

Reversible jump attack to textual classifiers with modification reduction

delete2024-04-22
delete0
delete
OA
AI
M
Mingze Ni
Z
Zhensu Sun
W
Wei Liu *
DOI:10.1007/s10994-024-06539-6delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recent studies on adversarial examples expose vulnerabilities of natural language processing models. Existing techniques for generating adversarial examples are typically driven by deterministic hierarchical rules that are agnostic to the optimal adversarial examples, a strategy that often results in adversarial samples with a suboptimal balance between magnitudes of changes and attack successes. To this end, in this research we propose two algorithms, Reversible Jump Attack (RJA) and Metropolis-Hasting Modification Reduction (MMR), to generate highly effective adversarial examples and to improve the imperceptibility of the examples, respectively. RJA utilizes a novel randomization mechanism to enlarge the search space and efficiently adapts to a number of perturbed words for adversarial examples. With these generated adversarial examples, MMR applies the Metropolis-Hasting sampler to enhance the imperceptibility of adversarial examples. Extensive experiments demonstrate that RJA-MMR outperforms current state-of-the-art methods in attack performance, imperceptibility, fluency and grammar correctness.
Keywords:
Textual attack
Adversarial learning
Natural language processing

Journal

Machine Learning cover
Machine Learning
IF:
2.9
Papers:
2.6K
Citations:
3.4W

Organization

U
university of technology sydney
Scholars:
1.6W
Papers: 2.0W
Citations: 25
S
ShanghaiTech University
Scholars:
9.6K
Papers: 5.9K
Citations: 1.6W