Return
Character-Level Adversarial Samples Generation Method Based on Dual-Layer Text Watermark
DOI:10.1109/ACCESS.2025.3594727.png)
Abstract
En 中文
Textual adversarial samples are widely used to assess the robustness and security of language models. Most existing methods generate these samples by substitution or deletion. However, such approaches are often easy to detect and show limited attack effectiveness. To address these issues, we propose a character-level adversarial attack framework based on dual-layer watermarking. The method embeds explicit and implicit watermarks in the original text to interfere with downstream models. First, we apply character-level gradient optimization to perturb explicit statistical features. Then, we use reinforcement learning to fine-tune semantic and encoding characteristics, guided by feedback from the target model. This approach ensures that the generated samples significantly reduce the performance of target models while preserving semantic meaning. Experiments on four benchmark datasets show that our method outperforms state-of-the-art baselines. It reduces average word distance (WD) by 12.5%, improves BERTScore by 0.9%, lowers the attack success rate (TMASR) by 11.0%, and increases the concealment rate (Hide) by 5.6%.
Keywords:
Large language models
adversarial attacks
adversarial sample generation
dual-layer watermark
text watermark
Journal
IF:
3.6
Papers:
9.8W
Citations:
29.4W

