1
Return

SEAttack: A self-evolving jailbreak attack to induce toxic responses for non-toxic queries in large language models

delete2025-12-12
delete0
PRE
AI
H
Huijun Liu
S
Shasha Li
B
Bin Ji *
X
Xiaohu Du
X
Xiaopeng Li
J
Jun null *
J
Jie Yu *
DOI:10.1016/j.ipm.2025.104544delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
To prevent Large Language Models (LLMs) from being misused by bad actors, model developers delve into designing safety mechanisms to guarantee LLMs generate helpful, honest, and harmless responses. To disclose flaws in the safety mechanisms, researchers conduct adversarial attacks against LLMs to circumvent the safety mechanisms, which are called “jailbreak attack”. Existing jailbreak attacks solely focus on engineering toxic queries and adversarial prompts to induce LLMs to generate toxic responses, losing sight of the study starting from engineering non-toxic queries and adversarial prompts, which also plays an important role in disclosing flaws in safety mechanisms. To fill the research gap, we propose SEAttack, a self-evolving jailbreak attack method to induce LLMs to generate toxic responses for non-toxic queries. Given a non-toxic query, SEAttack initially generates a response, which is more likely to be non-toxic due to the safety mechanisms. Then, it uses multiple iterations of self-evolving to evolve the non-toxic response to a toxic one. To evaluate SEAttack, we construct JailChat, a dataset containing 3000 non-toxic queries. We drive SEAttack to attack eighteen state-of-the-art LLMs, including five closed-source and thirteen open-source LLMs. Experimental results demonstrate that SEAttack achieves up to 89.07 % attack success rate, revealing non-negligible flaws in LLMs’ safety mechanisms. Moreover, we track the changes in the safety mechanisms of four ChatGPT variants. Extensive analyses and human evaluation further validate the effectiveness and rationality of SEAttack.

Journal

I
Information Processing and Management
IF:
6.9
Papers:
5.2K
Citations:
1.4W

Organization

N
National University of Defense Technology
Scholars:
3.3K
Papers: 1.0K
Citations: 8.2K
H
huazhong university of science and technology
Scholars:
2.3W
Papers: 7.2K
Citations: 5
Cited Papers

Cited Papers

Citing Papers

Citing Papers