Return
Multigranularity Adversarial Attacks on Large Language Models Using Genetic Programming
W
M
Y
Y
A
L
DOI:10.1109/tevc.2025.3629409.png)
Abstract
En 中文
large language models (LLMs) have demonstrated remarkable capabilities across various natural language processing tasks, but they remain vulnerable to adversarial attacks and pose significant security concerns. Existing attack methods often treat adversarial prompts as flat sequences, neglecting the rich hierarchical structure of natural language, which could limit their effectiveness. Advancing the methodologies for adversarial attacks is crucial for rigorously assessing the security of LLMs and identifying subtle vulnerabilities. This article introduces AdvGP, a novel framework that leverages genetic programming (GP) to generate adversarial prompts for LLMs. AdvGP exploits the inherent structural similarities between GP trees and natural language syntax to optimize the structure of harmful prompts. The framework incorporates a multigranularity hierarchical attack strategy, specialized genetic operators that leverage an assisting LLM for depth-aware crossover and multilevel mutation, and a comprehensive fitness function integrating semantic consistency and attack effectiveness. The proposed method achieves competitive attack performance on multiple LLMs, consistently generating harmful outputs despite higher perplexity than some baselines. Ablation studies confirm the significant contributions of both LLM-aided and depth-aware mechanisms to AdvGP’s effectiveness. Furthermore, transferability analysis reveals that the generated prompts are able to bypass the defenses of various state-of-the-art LLMs, such as ChatGPT and Gemini.
Keywords:
Adversarial attack
genetic programming (GP)
large language models (LLMs)
natural language processing
Journal
IF:
12
Papers:
1.8K
Citations:
2.4W
