Return
SEAttack: A self-evolving jailbreak attack to induce toxic responses for non-toxic queries in large language models
H
S
B
X
X
J
J
DOI:10.1016/j.ipm.2025.104544.png)
Abstract
En 中文
To prevent Large Language Models (LLMs) from being misused by bad actors, model developers delve into designing safety mechanisms to guarantee LLMs generate helpful, honest, and harmless responses. To disclose flaws in the safety mechanisms, researchers conduct adversarial attacks against LLMs to circumvent the safety mechanisms, which are called “jailbreak attack”. Existing jailbreak attacks solely focus on engineering toxic queries and adversarial prompts to induce LLMs to generate toxic responses, losing sight of the study starting from engineering non-toxic queries and adversarial prompts, which also plays an important role in disclosing flaws in safety mechanisms. To fill the research gap, we propose SEAttack, a self-evolving jailbreak attack method to induce LLMs to generate toxic responses for non-toxic queries. Given a non-toxic query, SEAttack initially generates a response, which is more likely to be non-toxic due to the safety mechanisms. Then, it uses multiple iterations of self-evolving to evolve the non-toxic response to a toxic one. To evaluate SEAttack, we construct JailChat, a dataset containing 3000 non-toxic queries. We drive SEAttack to attack eighteen state-of-the-art LLMs, including five closed-source and thirteen open-source LLMs. Experimental results demonstrate that SEAttack achieves up to 89.07 % attack success rate, revealing non-negligible flaws in LLMs’ safety mechanisms. Moreover, we track the changes in the safety mechanisms of four ChatGPT variants. Extensive analyses and human evaluation further validate the effectiveness and rationality of SEAttack.
Journal
I
IF:
6.9
Papers:
5.2K
Citations:
1.4W
