arrow
Return

Character-Level Adversarial Samples Generation Method Based on Dual-Layer Text Watermark

delete2025-01-01
delete0
delete
OA
AI
H
Houyue Wu
X
Xianwei Li
S
Shunxiang Zhang
H
Honghao Zhu
DOI:10.1109/ACCESS.2025.3594727delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Textual adversarial samples are widely used to assess the robustness and security of language models. Most existing methods generate these samples by substitution or deletion. However, such approaches are often easy to detect and show limited attack effectiveness. To address these issues, we propose a character-level adversarial attack framework based on dual-layer watermarking. The method embeds explicit and implicit watermarks in the original text to interfere with downstream models. First, we apply character-level gradient optimization to perturb explicit statistical features. Then, we use reinforcement learning to fine-tune semantic and encoding characteristics, guided by feedback from the target model. This approach ensures that the generated samples significantly reduce the performance of target models while preserving semantic meaning. Experiments on four benchmark datasets show that our method outperforms state-of-the-art baselines. It reduces average word distance (WD) by 12.5%, improves BERTScore by 0.9%, lowers the attack success rate (TMASR) by 11.0%, and increases the concealment rate (Hide) by 5.6%.
Keywords:
Large language models
adversarial attacks
adversarial sample generation
dual-layer watermark
text watermark

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

A
Anhui University of Science and Technology
Scholars:
2.1K
Papers: 706
Citations: 6.3K
B
Bengbu University
Scholars:
558
Papers: 334
Citations: 181