Skip to main content
QUICK REVIEW

[Paper Review] Sparks of Artificial General Intelligence: Early experiments with GPT-4

Sébastien Bubeck, Varun Chandrasekaran|arXiv (Cornell University)|Mar 22, 2023
Artificial Intelligence in Healthcare and Education1,528 citations
TL;DR

The paper presents an early study of GPT-4, arguing it exhibits broad, human-level capabilities across language, mathematics, coding, vision, medicine, law, and more, suggesting it is a step toward AGI while noting limitations and societal implications.

ABSTRACT

Artificial intelligence (AI) researchers have been developing and refining large language models (LLMs) that exhibit remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. The latest model developed by OpenAI, GPT-4, was trained using an unprecedented scale of compute and data. In this paper, we report on our investigation of an early version of GPT-4, when it was still in active development by OpenAI. We contend that (this early version of) GPT-4 is part of a new cohort of LLMs (along with ChatGPT and Google's PaLM for example) that exhibit more general intelligence than previous AI models. We discuss the rising capabilities and implications of these models. We demonstrate that, beyond its mastery of language, GPT-4 can solve novel and difficult tasks that span mathematics, coding, vision, medicine, law, psychology and more, without needing any special prompting. Moreover, in all of these tasks, GPT-4's performance is strikingly close to human-level performance, and often vastly surpasses prior models such as ChatGPT. Given the breadth and depth of GPT-4's capabilities, we believe that it could reasonably be viewed as an early (yet still incomplete) version of an artificial general intelligence (AGI) system. In our exploration of GPT-4, we put special emphasis on discovering its limitations, and we discuss the challenges ahead for advancing towards deeper and more comprehensive versions of AGI, including the possible need for pursuing a new paradigm that moves beyond next-word prediction. We conclude with reflections on societal influences of the recent technological leap and future research directions.

Motivation & Objective

  • Demonstrate GPT-4's broad, cross-domain capabilities beyond language alone.
  • Assess whether GPT-4 exhibits general intelligence or emergent behaviors close to human performance.
  • Investigate GPT-4's limitations, failure modes, and biases to outline challenges on the path to AGI.
  • Discuss the societal influences and governance considerations of a potential general AI leap.

Proposed method

  • Interact with an early GPT-4 instance using natural language prompts across diverse domains (language, math, coding, vision, medicine, law, psychology).
  • Compare GPT-4 outputs with prior models (e.g., ChatGPT) to assess generality and performance gaps.
  • Elicit targeted tasks (e.g., multimodal reasoning, tool use, planning) to probe general-purpose capabilities beyond memorization.
  • Vary prompts to test adaptability, stylistic flexibility, and problem-solving approaches.
  • Document limitations, biases, and failure modes to identify barriers to deeper AGI capabilities.

Experimental results

Research questions

  • RQ1Does GPT-4 demonstrate general, cross-domain abilities beyond language tasks?
  • RQ2To what extent does GPT-4 approach human-level performance across diverse domains without task-specific prompting?
  • RQ3What are the primary limitations, failure modes, and biases that constrain GPT-4’s general intelligence?
  • RQ4What societal and ethical implications accompany a system exhibiting broad, AGI-like capabilities?

Key findings

  • GPT-4 exhibits capabilities across mathematics, coding, vision, medicine, law, and psychology in addition to language.
  • GPT-4’s performance in many tasks is close to human-level and often surpasses prior models like ChatGPT.
  • GPT-4 demonstrates emergent, non-human-like patterns of intelligence and adaptability across domains.
  • The model shows limitations in planning, arithmetic, and some reasoning tasks, highlighting gaps toward full AGI.
  • There are notable concerns about misinformation, bias, and societal impact that accompany advanced LLM capabilities.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.