[논문 리뷰] OpenAI GPT-5 System Card
GPT-5는 실시간 라우터가 있는 빠른 모델 계층과 사고 모델 계층을 도입하여 안전성(안전한 완성), 환각 감소, 건강, 코딩, 다국어 작업 전반의 성능 향상을 이루고, 광범위한 레드-팀 점검과 안전장치를 갖춘 시스템을 제공합니다.
This is the system card published alongside the OpenAI GPT-5 launch, August 2025. GPT-5 is a unified system with a smart and fast model that answers most questions, a deeper reasoning model for harder problems, and a real-time router that quickly decides which model to use based on conversation type, complexity, tool needs, and explicit intent (for example, if you say 'think hard about this' in the prompt). The router is continuously trained on real signals, including when users switch models, preference rates for responses, and measured correctness, improving over time. Once usage limits are reached, a mini version of each model handles remaining queries. This system card focuses primarily on gpt-5-thinking and gpt-5-main, while evaluations for other models are available in the appendix. The GPT-5 system not only outperforms previous models on benchmarks and answers questions more quickly, but -- more importantly -- is more useful for real-world queries. We've made significant advances in reducing hallucinations, improving instruction following, and minimizing sycophancy, and have leveled up GPT-5's performance in three of ChatGPT's most common uses: writing, coding, and health. All of the GPT-5 models additionally feature safe-completions, our latest approach to safety training to prevent disallowed content. Similarly to ChatGPT agent, we have decided to treat gpt-5-thinking as High capability in the Biological and Chemical domain under our Preparedness Framework, activating the associated safeguards. While we do not have definitive evidence that this model could meaningfully help a novice to create severe biological harm -- our defined threshold for High capability -- we have chosen to take a precautionary approach.
연구 동기 및 목표
- GPT-5를 빠른 모델과 사고 모델 및 실시간 라우터를 갖춘 통합 시스템으로 도입합니다.
- 데이터, 학습 및 안전 지향 방법론(안전한 완성, 거절 조정, 및 안전장치)을 설명합니다.
- 안전 문제, 환각, 기만, 탈옥, 다국어 성능을 이전 모델과 비교하여 평가합니다.
- 고위험 영역(bio/chem, 사이버 보안)에 대한 Preparedness Framework의 레드-팀 활동, 외부 평가 및 준비체계를 상세히 제공합니다.
- 거짓 말하기에 대한 관리, 순응성 감소 및 안전성 향상을 위한 거버넌스, 모니터링 및 향후 방향을 개괄합니다.]
- method_1번문장은
- method_2번문장은
- method_3번문장은
제안 방법
- 모델 분류 체계를 정의합니다: gpt-5-main, gpt-5-main-mini (빠른 모델) 및 gpt-5-thinking, gpt-5-thinking-mini, gpt-5-thinking-nano (사고 모델).
- 합리화 모델을 강화학습으로 학습하여 이진 거절보다 안전한 완성에 초점을 맞춥니다.
- System, Developer, User 메시지를 포함하는 지시 계층(Instructions Hierarchy)을 통해 다층 방어 스택을 사용합니다.
- 금지된 콘텐츠, 탈옥 강건성, 프롬프트 주입, 환각 벤치마크(LongFact, FActScore, HealthBench) 등을 통해 안전성을 평가합니다.
- 폭력적 공격 계획 및 프롬프트 주입 테스트를 포함한 광범위한 레드-팀 점검(5000+ 시간, 400+ 테스터) 및 외부 연구자/벤더와의 평가를 진행합니다.
- HealthBench, MMLU 다국어 벤치마킹, BBQ 공정성 평가, 안전성/성능 비교를 GPT-4o 및 OpenAI o3 대비 적용합니다.

실험 결과
연구 질문
- RQ1안전한 완성이 전통적인 거절 대비 안전 실패를 줄이고 도움성을 증가시키는지 어떻게 비교됩니다?
- RQ2실세계 작업(건강, 코딩, 다국어)에서 gpt-5-main과 gpt-5-thinking 간의 안전성과 성능 트레이드오프는 무엇인가요?
- RQ3 jailbreak, 프롬프트 주입 및 지시 계층 완화가 GPT-5 모델들의 안전성에 어떤 영향을 주나요?
- RQ4안전장치가 추론 작업에서 기만, 순응성, 환각에 미치는 영향은 무엇인가요?
- RQ5외부 레드-팀 평가가 시스템 차원의 취약점 식별에 어떤 차이를 보이나요?
주요 결과
- gpt-5-thinking 및 gpt-5-main은 이전 모델 대비 안전성과 도움성이 향상되었고, 환각 및 순응성이 감소했습니다.
- 표준 금지 콘텐츠 지표에서 모든 모델이 높은 안전성을 보이고 있으며, 생산 벤치마크에서 미세한 개선과 혐오/괴롭힘 범주에서의 일부 후퇴가 확인됩니다.
- 환각률이 크게 감소합니다: gpt-5-main의 사실 오류는 GPT-4o보다 26% 낮고, gpt-5-thinking의 경우 OpenAI o3보다 65% 낮습니다.
- 온라인 및 오프라인 평가에서 순응성이 현저히 감소했고, GPT-4o 대비 큰 감소를 보였습니다.
- HealthBench 결과에서 gpt-5-thinking은 이전 모델 대비 건강 안전성 및 성능이 크게 향상되었으며, 환각 및 긴급 상황 오류가 크게 감소했습니다.
- chain-of-thought를 통한 기만 모니터링에서 gpt-5-thinking이 o3 대비 기만 비율이 더 낮았습니다(약 2.1% 대 약 4.8%).
- 이미지 입력 안전성 및 다국어 능력(13-language MMLU)이 기준선 대비 경쟁력 있는 성능을 보입니다.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.