1
Return

RogueGPT: Unleashing Jailbreak Prompts on LLMs A Comparative Analysis of Security Efficiency Across Large Language Models

delete2026-04-01
delete0
PRE
AI
S
Shivaswaroopa, Arpitha
S
Sood, Vanshika
G
Gururaj, H. L.
S
Shreyas, J.
A
Al-Turjman, Fadi *
DOI:10.1002/eng2.70069delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Large Language Models (LLMs) have seen a remarkable surge in popularity since the latter part of 2022. These models have become vital in the lives of individuals from varying professions. While some users leverage LLMs for academic or informational purposes, others exploit them for illicit activities. Methods of exploitation include Adversarial Attacks, Instruction Tuning Attacks, Inference Attacks, and Extraction Attacks. This paper investigates a specific Instruction Tuning Attack known as jailbreaking, which manipulates LLMs with prompts to generate harmful responses to forbidden instructions. This study presents compelling evidence of how widely used LLMs, such as OpenAI's ChatGPT, Google's Gemini, Meta's LLaMa, LMSYS's Vicuna, and Alibaba Cloud's Qwen, can be manipulated to generate responses that range from mildly illegal to potentially criminal content. Jailbreak prompts were created for each LLM, encompassing a range of inquiries spanning various categories. Based on the level of response elicited, they were categorized and computed alongside the Attack-to-Success Rate (ASR). These findings highlight the effectiveness of our prompts on each LLM and their performance relative to other models. Vicuna produced the best results with ASR (0.93) and FT (0.842), followed by LLaMa with ASR (0.71) and FT (0.709), indicating their vulnerability. The category of False Information had the highest overall average, with ASR (0.864) and FT (0.96). Our conclusions were reached through a combination of human assessment and quantitative analysis, detailed in subsequent sections. Through the dissemination of this research, the aim is to encourage organizations to prioritize their security measures and raise awareness among individuals about the responsible and ethical use of LLMs, given their potential for harm.
Keywords:
Attack-to-Success Rate
Fine-Tuning
Generative Artificial Intelligence (Gen AI)
jailbreak prompt
Large Language Models (LLMs)
RogueGPT

Journal

Engineering Reports cover
Engineering Reports
IF:
2
Papers:
362
Citations:
1.7K

Organization

Near East University cover
Near East University
Scholars:
411
Papers: 297
Citations: 2.3K
M
manipal academy of higher education (mahe)
Scholars:
1.7K
Papers: 644
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers