[论文解读] Coordinated pausing: An evaluation-based coordination scheme for frontier AI developers
本文提出了一种基于评估的协调机制,供前沿人工智能开发者在检测到大规模模型中存在危险能力时集体暂停研究。该机制包括标准化评估、失败后强制暂停、跨开发者通知、安全分析,并仅在满足安全阈值后才恢复研究——为新兴人工智能风险提供一种结构化、可扩展的应对方案。
As artificial intelligence (AI) models are scaled up, new capabilities can emerge unintentionally and unpredictably, some of which might be dangerous. In response, dangerous capabilities evaluations have emerged as a new risk assessment tool. But what should frontier AI developers do if sufficiently dangerous capabilities are in fact discovered? This paper focuses on one possible response: coordinated pausing. It proposes an evaluation-based coordination scheme that consists of five main steps: (1) Frontier AI models are evaluated for dangerous capabilities. (2) Whenever, and each time, a model fails a set of evaluations, the developer pauses certain research and development activities. (3) Other developers are notified whenever a model with dangerous capabilities has been discovered. They also pause related research and development activities. (4) The discovered capabilities are analyzed and adequate safety precautions are put in place. (5) Developers only resume their paused activities if certain safety thresholds are reached. The paper also discusses four concrete versions of that scheme. In the first version, pausing is completely voluntary and relies on public pressure on developers. In the second version, participating developers collectively agree to pause under certain conditions. In the third version, a single auditor evaluates models of multiple developers who agree to pause if any model fails a set of evaluations. In the fourth version, developers are legally required to run evaluations and pause if dangerous capabilities are discovered. Finally, the paper discusses the desirability and feasibility of our proposed coordination scheme. It concludes that coordinated pausing is a promising mechanism for tackling emerging risks from frontier AI models. However, a number of practical and legal obstacles need to be overcome, especially how to avoid violations of antitrust law.
研究动机与目标
- 解决前沿人工智能模型在扩展过程中出现危险能力时缺乏系统性响应协议的问题。
- 回应日益增长的担忧:即扩展人工智能模型可能意外触发高风险行为,如操纵、网络攻击或自主自我复制。
- 开发一种可行且可扩展的协调机制,使多个开发者能在检测到危险能力时协同行动。
- 探索实施协调暂停的实用与法律路径,同时避免违反反垄断法。
- 提供一种将安全评估整合到人工智能开发生命周期的框架,明确暂停与恢复研究的触发条件。
提出的方法
- 实施五步基于评估的协调机制:(1) 评估前沿模型的危险能力,(2) 若任一模型失败则暂停开发,(3) 通知其他开发者,(4) 分析风险并实施安全措施,(5) 仅在满足安全阈值后才恢复。
- 提出该机制的四种版本:自愿公众压力、集体协议、第三方审计和法律强制合规。
- 以现有的危险能力评估(如对权力寻求行为的评估)为基础,同时呼吁开发新的标准化评估。
- 整合安全阈值,以风险评估为基础,明确何时触发暂停以及何时允许恢复。
- 通过探讨现有法律中的潜在安全港条款(如美国《国防生产法》第708条),解决法律可行性问题。
- 鼓励各实验室内部制定响应政策,参考 Anthropic 的负责任扩展政策和 ARC Evals 的框架。

实验结果
研究问题
- RQ1当某一开发者在模型中发现危险能力时,前沿人工智能开发者如何协调暂停研究?
- RQ2实现协调暂停的最可行且理想机构形式是什么?
- RQ3协调暂停面临的主要法律与实际障碍(尤其是反垄断担忧)是什么?
- RQ4如何开发并验证可靠、标准化的危险能力评估?
- RQ5应设定何种安全阈值以决定何时暂停及何时恢复研究?
主要发现
- 协调暂停是缓解前沿人工智能模型潜在新兴风险的有前景机制,尤其在检测到自我复制或自主目标追求等能力时。
- 在受访专家中,98% 的人认为人工智能实验室应运行危险能力评估,93% 支持在发现此类风险时暂停研究。
- 该提议的机制可通过四种不同模式实施:自愿公众压力、集体协议、第三方审计或法律强制。
- 现有的危险能力评估(如对权力寻求行为的评估)虽有限,但可作为进一步发展的基础。
- 法律问题,尤其是反垄断法方面的担忧,仍是重大障碍,但潜在的安全港条款(如美国《国防生产法》第708条)或可提供解决路径。
- 该框架呼吁集体协调与各实验室内部响应政策并行实施,基于评估结果设定明确的暂停与恢复触发条件。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。