[Paper Review] Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
The study classifies jailbreak prompts into a taxonomy, empirically tests their ability to bypass ChatGPT restrictions across GPT-3.5-TURBO and GPT-4 using 3,120 jailbreak questions in eight prohibited scenarios, and analyzes model robustness and prompt evolution.
Large Language Models (LLMs), like ChatGPT, have demonstrated vast potential but also introduce challenges related to content constraints and potential misuse. Our study investigates three key research questions: (1) the number of different prompt types that can jailbreak LLMs, (2) the effectiveness of jailbreak prompts in circumventing LLM constraints, and (3) the resilience of ChatGPT against these jailbreak prompts. Initially, we develop a classification model to analyze the distribution of existing prompts, identifying ten distinct patterns and three categories of jailbreak prompts. Subsequently, we assess the jailbreak capability of prompts with ChatGPT versions 3.5 and 4.0, utilizing a dataset of 3,120 jailbreak questions across eight prohibited scenarios. Finally, we evaluate the resistance of ChatGPT against jailbreak prompts, finding that the prompts can consistently evade the restrictions in 40 use-case scenarios. The study underscores the importance of prompt structures in jailbreaking LLMs and discusses the challenges of robust jailbreak prompt generation and prevention.
Motivation & Objective
- Identify and classify jailbreak prompt types and patterns.
- Quantify jailbreak effectiveness across prohibited scenarios and model versions.
- Assess robustness and evolution of jailbreak prompts over time.
- Examine factors influencing protection strength in different GPT models and policies.
Proposed method
- Collected 78 verified jailbreak prompts from jailbreak chat sources up to April 2023.
- Developed a jailbreak prompt categorization model identifying 10 patterns within 3 types (pretending, attention shifting, privilege escalation).
- Created 40 scenario prompts aligned with OpenAI disallowed usage policy across 8 prohibited scenarios.
- Conducted 31,200 queries (5 rounds × 8 scenarios × 78 prompts × 2 models) on GPT-3.5-TURBO and GPT-4.
- Manually evaluated whether responses violated prohibitions, and analyzed prompt evolution and defense gaps.
Experimental results
Research questions
- RQ1RQ1: How many types and patterns of jailbreak prompts exist and how are they distributed?
- RQ2RQ2: How capable are jailbreak prompts at bypassing LLM restrictions across scenarios and model versions?
- RQ3RQ3: How strong is CHATGPT’s protection against jailbreak prompts, and how does it vary by model version and policy?
Key findings
- Pretending is the dominant jailbreak strategy (97.44% of prompts).
- Easiest prohibited scenarios to jailbreak are Illegal Activities (IA), Fraudulent/Deceptive Activities (FDA), and Adult Content (ADULT).
- Simulate Jailbreaking (SIMU) and Superior Model (SUPER) are the most effective patterns (≈93% success).
- Program Execution (PROG) is the least effective pattern (≈69% success).
- GPT-4 reduces jailbreak success rates relative to GPT-3.5-TURBO by about 15.5% on average, with larger reductions for Harmful Content (≈38.4%).
- DAN-style prompt evolution shows increasing jailbreak success over time, indicating ongoing adaptation by adversaries.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.