Skip to main content
QUICK REVIEW

[论文解读] Open Problems in Cooperative AI

Allan Dafoe, Edward Hughes|arXiv (Cornell University)|Dec 15, 2020
Evolutionary Game Theory and Cooperation参考文献 322被引用 17
一句话总结

本文提出将协作型人工智能(Cooperative AI)作为一项新的研究方向,聚焦于赋予人工智能系统与人类及其他智能体有效协作所需的能力。文章概述了关键的协作能力——理解、沟通、承诺与制度,并强调跨学科整合以及减轻胁迫和排斥等风险。

ABSTRACT

Problems of cooperation--in which agents seek ways to jointly improve their welfare--are ubiquitous and important. They can be found at scales ranging from our daily routines--such as driving on highways, scheduling meetings, and working collaboratively--to our global challenges--such as peace, commerce, and pandemic preparedness. Arguably, the success of the human species is rooted in our ability to cooperate. Since machines powered by artificial intelligence are playing an ever greater role in our lives, it will be important to equip them with the capabilities necessary to cooperate and to foster cooperation. We see an opportunity for the field of artificial intelligence to explicitly focus effort on this class of problems, which we term Cooperative AI. The objective of this research would be to study the many aspects of the problems of cooperation and to innovate in AI to contribute to solving these problems. Central goals include building machine agents with the capabilities needed for cooperation, building tools to foster cooperation in populations of (machine and/or human) agents, and otherwise conducting AI research for insight relevant to problems of cooperation. This research integrates ongoing work on multi-agent systems, game theory and social choice, human-machine interaction and alignment, natural-language processing, and the construction of social tools and platforms. However, Cooperative AI is not the union of these existing areas, but rather an independent bet about the productivity of specific kinds of conversations that involve these and other areas. We see opportunity to more explicitly focus on the problem of cooperation, to construct unified theory and vocabulary, and to build bridges with adjacent communities working on cooperation, including in the natural, social, and behavioural sciences.

研究动机与目标

  • 应对现实场景中人工智能系统与人类及其他智能体有效协作的日益增长的需求。
  • 识别并系统化协作所需的核心能力,如理解、沟通、承诺与制度设计。
  • 通过构建统一的理论与术语框架,弥合人工智能研究与社会科学之间的鸿沟,以应对协作问题。
  • 研究协作型人工智能的潜在风险,包括排斥、串通与胁迫,并引导其发展以促进人类福祉。
  • 推动人工智能、博弈论、社会选择理论与人机交互等领域之间的跨学科合作,以实现可扩展的协作解决方案。

提出的方法

  • 沿战略情境、共同利益与冲突利益、个体与规划者视角等维度对协作机会进行分类。
  • 将协作能力结构化为四大核心组件:理解(对世界、行为、偏好及递归信念的理解)、沟通(共同基础、带宽、教学、混合动机)、承诺(单边/多边、条件/无条件)以及制度(去中心化规范、集中化系统、信任机制)。
  • 提出在机器学习中开发训练环境与任务,使协作技能成为必要、可学习且非平凡的。
  • 整合多智能体系统、博弈论、机制设计、自然语言处理与可解释性等领域的洞见,以构建协作型人工智能系统。
  • 通过偏好学习与对齐技术,将人类价值观与伦理规范融入协作型人工智能。
  • 设计支持人类与机器智能体群体协作的工具与平台,包括声誉系统与调解算法。

实验结果

研究问题

  • RQ1如何设计人工智能智能体,使其在协作环境中能够理解其他智能体的信念、偏好与行为?
  • RQ2在混合动机条件下,何种沟通机制能够实现高效、稳健且公平的协作?
  • RQ3何种形式的协作承诺——单边、多边、条件性或硬件嵌入式——最能支持长期协调?
  • RQ4如何设计制度与规范,以支持人工智能与人类智能体之间可扩展的去中心化协作?
  • RQ5协作型人工智能能力可能带来的风险(如排斥或胁迫)是什么?如何通过技术和制度设计加以缓解?

主要发现

  • 协作型人工智能并非现有人工智能子领域的简单整合,而是一项需要专注且跨学科对话的独立研究议程。
  • 本文识别出四大核心能力——理解、沟通、承诺与制度——是构建协作型人工智能系统所不可或缺的。
  • 人工智能中的有效协作需要整合博弈论、社会选择、多智能体系统与人机交互的洞见。
  • 协作型人工智能的发展必须关注如串通、胁迫与排斥等风险,尤其是在高自主性系统中。
  • 机器学习中的训练环境与任务应被设计为使协作技能既关键又可学习,从而实现可扩展的协作。
  • 该研究议程呼吁建立统一的理论与共享术语,以连接人工智能研究与更广泛的科学与社会协作努力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。