Skip to main content
QUICK REVIEW

[Paper Review] Levels of AGI for Operationalizing Progress on the Path to AGI

Meredith Ringel Morris, Jascha Sohl‐Dickstein|arXiv (Cornell University)|Nov 4, 2023
Data Quality and Management45 citations
TL;DR

The paper proposes a two-dimensional, leveled ontology for AGI based on performance (depth) and generality (breadth), defines six guiding principles, and discusses benchmarks, risk, and human–AI interaction implications along the path to AGI.

ABSTRACT

We propose a framework for classifying the capabilities and behavior of Artificial General Intelligence (AGI) models and their precursors. This framework introduces levels of AGI performance, generality, and autonomy, providing a common language to compare models, assess risks, and measure progress along the path to AGI. To develop our framework, we analyze existing definitions of AGI, and distill six principles that a useful ontology for AGI should satisfy. With these principles in mind, we propose "Levels of AGI" based on depth (performance) and breadth (generality) of capabilities, and reflect on how current systems fit into this ontology. We discuss the challenging requirements for future benchmarks that quantify the behavior and capabilities of AGI models against these levels. Finally, we discuss how these levels of AGI interact with deployment considerations such as autonomy and risk, and emphasize the importance of carefully selecting Human-AI Interaction paradigms for responsible and safe deployment of highly capable AI systems.

Motivation & Objective

  • Clarify a clear, operational definition of AGI focused on capabilities, generality, and performance.
  • Provide a leveled taxonomy (Levels of AGI) to track progress along the path to AGI.
  • Outline principles for measuring AGI via benchmarks and ecologically valid tasks.
  • Discuss risk, autonomy, and Human-AI Interaction considerations at different levels.
  • Suggest how deployment and interaction paradigms influence safe use of capable AI systems.

Proposed method

  • Develop a two-dimensional leveling framework (depth of performance x breadth of generality) for AGI.
  • Derive six guiding principles for a useful AGI ontology (capabilities, generality, cognitive/metacognitive tasks, potential vs deployment, ecological validity, path vs endpoint).
  • Propose a matrixed table mapping levels for different tasks and systems (e.g., Emerging, Competent, Expert, Virtuoso, ASI).
  • Discuss benchmark design considerations and the notion of a living benchmark with task-generation capabilities.
  • Analyze risk contexts (autonomy, interface, governance) and relate levels to potential risks.
  • Describe Human-AI Interaction levels of autonomy and related deployment considerations.

Experimental results

Research questions

  • RQ1How should AGI be defined in a way that emphasizes capabilities, generality, and autonomy rather than underlying mechanisms?
  • RQ2What levels of performance and generality best capture progress toward AGI, and how can these be measured?
  • RQ3What benchmarks and task sets would meaningfully assess progression along the path to AGI?
  • RQ4How do levels of AGI interact with deployment considerations and risk, including autonomy and human-AI interaction paradigms?

Key findings

  • A two-dimensional Levels of AGI framework (depth of performance x breadth of generality) is proposed to classify systems along the path to AGI.
  • A set of six principles guides a useful AGI ontology, emphasizing capabilities, generality, cognitive/metacognitive tasks, potential over deployment, ecological validity, and progression along the path to AGI.
  • Current frontier models may span multiple levels (e.g., Emerging AGI in some tasks, Competent or higher in others), highlighting the need for ecologically valid benchmarks and model documentation.
  • Benchmarking toward AGI should be a living process with open-ended tasks and a framework to add new tasks as capabilities evolve.
  • The framework discusses how different levels relate to deployment considerations, autonomy, and risk, arguing for a nuanced, not endpoint-focused, view of AGI progression.
  • The paper connects levels to existing definitions (e.g., OpenAI’s labor substitution threshold) and emphasizes risk at higher levels (misalignment, autonomy risks).

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.