[Paper Review] GPT-4o System Card
OpenAI presents GPT-4o System Card detailing the multimodal GPT-4o model, its capabilities, safety evaluations, risk mitigations, and third-party assessments.
GPT-4o is an autoregressive omni model that accepts as input any combination of text, audio, image, and video, and generates any combination of text, audio, and image outputs. It's trained end-to-end across text, vision, and audio, meaning all inputs and outputs are processed by the same neural network. GPT-4o can respond to audio inputs in as little as 232 milliseconds, with an average of 320 milliseconds, which is similar to human response time in conversation. It matches GPT-4 Turbo performance on text in English and code, with significant improvement on text in non-English languages, while also being much faster and 50\% cheaper in the API. GPT-4o is especially better at vision and audio understanding compared to existing models. In line with our commitment to building AI safely and consistent with our voluntary commitments to the White House, we are sharing the GPT-4o System Card, which includes our Preparedness Framework evaluations. In this System Card, we provide a detailed look at GPT-4o's capabilities, limitations, and safety evaluations across multiple categories, focusing on speech-to-speech while also evaluating text and image capabilities, and measures we've implemented to ensure the model is safe and aligned. We also include third-party assessments on dangerous capabilities, as well as discussion of potential societal impacts of GPT-4o's text and vision capabilities.
Motivation & Objective
- Present GPT-4o capabilities across text, vision, and audio modalities and compare with GPT-4 Turbo in performance and cost.
- Describe data sources, pre-training, and data filtering/masking strategies used to mitigate risks.
- Outline the Preparedness Framework evaluations and safety mitigations across multiple risk categories.
- Detail external red-teaming processes, methodologies, and limitations of evaluations.
- Summarize third-party assessments and societal impact considerations for GPT-4o.
Proposed method
- Describe GPT-4o as an autoregressive omni model that processes text, image, audio, and video inputs and outputs text, audio, or image
- Explain data sources and training components including web data, code/math, and multimodal data
- Outline post-training alignment, red-teaming, and product-level mitigations as safety measures
- Discuss evaluation methodology using red-team data and conversions of text-based tasks to audio-based tasks with TTS
- Present Preparedness Framework evaluations and how high-risk categories influence deployment decisions
- Summarize third-party assessments (METR and Apollo Research) and their implications

Experimental results
Research questions
- RQ1What are GPT-4o's capabilities across text, vision, and audio modalities?
- RQ2How effective are the safety mitigations and moderation tools for speech-to-speech use cases?
- RQ3How does GPT-4o perform on diverse voices and accents in both capabilities and safety behavior?
- RQ4What are the outcomes of external red-teaming and third-party assessments in relation to autonomy-related risks?
Key findings
- GPT-4o matches GPT-4 Turbo on text in English and code and is faster and 50% cheaper in the API, with significant improvements in non-English text
- External red-teaming covered 45 languages across 29 countries and informed multiple safety evaluations and mitigations
- Voice mode mitigations show high accuracy in preventing unauthorized voice generation and speaker identification refusals (e.g., over 98% for should refuse in speaker identification)
- Disparate performance across diverse voices is marginal; safety behaviors are largely invariant across voices in evaluations
- Preparedness Framework classifies GPT-4o overall risk as medium after evaluating cybersecurity, CBRN, persuasion, and model autonomy
- Third-party assessments (METR and Apollo) provide additional validation but also highlight areas where autonomy-related capabilities are limited

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.