arrow
Return

Adaptive reinforcement learning based projected gradient descent attack

delete2026-07-04
delete0
PRE
AI
Z
Zihan Zhu
Y
Yuexin Zhang *
A
Ayong Ye
X
Xiaoding Wang
C
Chengling Wang
T
Tianqing Zhu *
DOI:10.1007/s11227-026-08687-zdelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Large language models (LLMs) are widely deployed in intelligent systems but remain vulnerable to adversarial attacks, such as Greedy coordinate gradient (GCG) and Projected gradient descent (PGD). Existing approaches suffer from three critical issues: (1) Gradient methods like PGD use static entropy factors, failing to adapt to the dynamic gap between continuous optimization and discrete evaluation, leading to low success rates; (2) Discrete methods like GCG rely on trial-and-error with 512 candidate tokens, incurring high computational costs; (3) Existing RL attacks optimize only the objective function, lacking component collaboration, which hinders breakthroughs against complex defenses. To address these, we propose Adaptive reinforcement learning-based projected gradient descent (ARL-PGD). It features three logically sequential innovations: (1) A reinforcement learning-based system-level framework that provides discrete evaluations as the system foundation; (2) A distributed discrete loss feedback mechanism to align continuous optimization with discrete objectives, mitigating GCG’s costs and PGD’s feedback gaps; (3)A dynamic entropy factor strategy: adapting entropy via relaxation gaps, it deterministically modulates distribution sharpness (distinct from noise uncertainty) to preserve optimization gains. These components form a closed-loop “evaluate-feedback-adjust” system, enabling nonlinear synergistic optimization and significantly improving attack success rates (ASR). Experiments on mainstream LLM models (Vicuna, Llama, and Gemma series) show ARL-PGD achieves higher ASR than baselines, with more natural and stealthy adversarial prompts. Ablation studies confirm each component’s effectiveness.
Keywords:
Adversarial attack
Large language model
Reinforcement learning
Projected gradient descent

Journal

T
The Journal of Supercomputing
IF:
0
Papers:
647
Citations:
0

Organization

F
Faculty of Data Science
Scholars:
43
Papers: 38
Citations: 0
I
industrial school of joint innovation
Scholars:
2
Papers: 1
Citations: 0
C
college of computer and cyber security
Scholars:
40
Papers: 16
Citations: 0
researcher View more organizations