[논문 리뷰] Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
연구는 jailbreak 프롬프트를 분류 체계로 나누고, 3,120개의 jailbreak 질문을 8가지 금지된 시나리오에서 GPT-3.5-TURBO와 GPT-4를 가로질러 채팅 시스템의 제한 우회를 평가하며, 모델의 견고성과 프롬프트의 진화를 분석한다.
Large Language Models (LLMs), like ChatGPT, have demonstrated vast potential but also introduce challenges related to content constraints and potential misuse. Our study investigates three key research questions: (1) the number of different prompt types that can jailbreak LLMs, (2) the effectiveness of jailbreak prompts in circumventing LLM constraints, and (3) the resilience of ChatGPT against these jailbreak prompts. Initially, we develop a classification model to analyze the distribution of existing prompts, identifying ten distinct patterns and three categories of jailbreak prompts. Subsequently, we assess the jailbreak capability of prompts with ChatGPT versions 3.5 and 4.0, utilizing a dataset of 3,120 jailbreak questions across eight prohibited scenarios. Finally, we evaluate the resistance of ChatGPT against jailbreak prompts, finding that the prompts can consistently evade the restrictions in 40 use-case scenarios. The study underscores the importance of prompt structures in jailbreaking LLMs and discusses the challenges of robust jailbreak prompt generation and prevention.
연구 동기 및 목표
- Identify and classify jailbreak prompt types and patterns.
- Quantify jailbreak effectiveness across prohibited scenarios and model versions.
- Assess robustness and evolution of jailbreak prompts over time.
- Examine factors influencing protection strength in different GPT models and policies.
제안 방법
- Collected 78 verified jailbreak prompts from jailbreak chat sources up to April 2023.
- Developed a jailbreak prompt categorization model identifying 10 patterns within 3 types (pretending, attention shifting, privilege escalation).
- Created 40 scenario prompts aligned with OpenAI disallowed usage policy across 8 prohibited scenarios.
- Conducted 31,200 queries (5 rounds × 8 scenarios × 78 prompts × 2 models) on GPT-3.5-TURBO and GPT-4.
- Manually evaluated whether responses violated prohibitions, and analyzed prompt evolution and defense gaps.]
- research_questions: ["RQ1: How many types and patterns of jailbreak prompts exist and how are they distributed?","RQ2: How capable are jailbreak prompts at bypassing LLM restrictions across scenarios and model versions?","RQ3: How strong is CHATGPT’s protection against jailbreak prompts, and how does it vary by model version and policy?"]
- key_findings:["Pretending is the dominant jailbreak strategy (97.44% of prompts).","Easiest prohibited scenarios to jailbreak are Illegal Activities (IA), Fraudulent/Deceptive Activities (FDA), and Adult Content (ADULT).","Simulate Jailbreaking (SIMU) and Superior Model (SUPER) are the most effective patterns (≈93% success).","Program Execution (PROG) is the least effective pattern (≈69% success).","GPT-4 reduces jailbreak success rates relative to GPT-3.5-TURBO by about 15.5% on average, with larger reductions for Harmful Content (≈38.4%).","DAN-style prompt evolution shows increasing jailbreak success over time, indicating ongoing adaptation by adversaries.
실험 결과
연구 질문
- RQ1RQ1: How many types and patterns of jailbreak prompts exist and how are they distributed?
- RQ2RQ2: How capable are jailbreak prompts at bypassing LLM restrictions across scenarios and model versions?
- RQ3RQ3: How strong is CHATGPT’s protection against jailbreak prompts, and how does it vary by model version and policy?
주요 결과
- Pretending is the dominant jailbreak strategy (97.44% of prompts).
- Easiest prohibited scenarios to jailbreak are Illegal Activities (IA), Fraudulent/Deceptive Activities (FDA), and Adult Content (ADULT).
- Simulate Jailbreaking (SIMU) and Superior Model (SUPER) are the most effective patterns (≈93% success).
- Program Execution (PROG) is the least effective pattern (≈69% success).
- GPT-4 reduces jailbreak success rates relative to GPT-3.5-TURBO by about 15.5% on average, with larger reductions for Harmful Content (≈38.4%).
- DAN-style prompt evolution shows increasing jailbreak success over time, indicating ongoing adaptation by adversaries.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.