[论文解读] Levels of AGI for Operationalizing Progress on the Path to AGI
论文提出一个基于表现(depth)和泛化广度(breadth)的二维分层AGI本体论,定义六条指导原则,并讨论基准、风险与在人机交互路径上的意义。
We propose a framework for classifying the capabilities and behavior of Artificial General Intelligence (AGI) models and their precursors. This framework introduces levels of AGI performance, generality, and autonomy, providing a common language to compare models, assess risks, and measure progress along the path to AGI. To develop our framework, we analyze existing definitions of AGI, and distill six principles that a useful ontology for AGI should satisfy. With these principles in mind, we propose "Levels of AGI" based on depth (performance) and breadth (generality) of capabilities, and reflect on how current systems fit into this ontology. We discuss the challenging requirements for future benchmarks that quantify the behavior and capabilities of AGI models against these levels. Finally, we discuss how these levels of AGI interact with deployment considerations such as autonomy and risk, and emphasize the importance of carefully selecting Human-AI Interaction paradigms for responsible and safe deployment of highly capable AI systems.
研究动机与目标
- 澄清一个清晰、可操作的AGI定义,聚焦能力、泛化能力和性能。
- 提供一个分层的分类法(AGI层级)以跟踪通向AGI的进展。
- 概述通过基准测试和符合生态有效性的任务来衡量AGI的原则。
- 讨论在不同层级上的风险、自治与人机交互考虑。
- 提出部署与交互范式如何影响对具备能力的AI系统的安全使用的建议。
提出的方法
- 为AGI开发一个二维分级框架(表现深度×泛化广度)。
- 推导六条有用的AGI本体论的指导原则(能力、泛化、认知/元认知任务、潜力与部署、生态有效性、路径与端点)。
- 提出一个矩阵表,为不同任务和系统映射水平(如Emerging、Competent、Expert、Virtuoso、ASI)。
- 讨论基准设计考虑因素以及具备任务生成能力的活基准的概念。
- 分析风险情境(自治、界面、治理)并将各水平与潜在风险相关联。
- 描述人机互动的自治水平及相关部署考虑。
实验结果
研究问题
- RQ1应以强调能力、泛化和自治而非底层机制的方式定义AGI吗?
- RQ2哪些表现和泛化水平最能体现朝向AGI的进展?如何测量?
- RQ3哪些基准和任务集能有意义地评估朝向AGI的进展?
- RQ4AGI的各层级如何与部署考量和风险(包括自治与人机互动范式)相互作用?
主要发现
- 提出了一个二维的AGI层级框架(表现深度×泛化广度)以沿着通向AGI的路径对系统进行分类。
- 六条原则引导有用的AGI本体论,强调能力、泛化、认知/元认知任务、潜力优先于部署、生态有效性以及通向AGI路径中的进展。
- 当前前沿模型在某些任务上可能跨越多个层级(如在某些任务上为Emerging AGI,在其他任务为Competent或更高),强调需要生态有效的基准和模型文档。
- 向AGI进行基准测试应是一个动态过程,具有开放式任务和随着能力演进添加新任务的框架。
- 该框架讨论了不同层级如何与部署考量、自治和风险相关,并主张对AGI进展采取细致入微、非端点导向的视角。
- 本文将层级与现有定义(如OpenAI的劳动替代阈值)联系起来,并强调在更高层级的风险(错配、自治风险)。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。