Skip to main content
QUICK REVIEW

[论文解读] Impossibility Results in AI: A Survey

Mario Brčić, Roman V. Yampolskiy|arXiv (Cornell University)|Sep 1, 2021
Ethics and Social Impacts of AI被引用 4
一句话总结

本综述将人工智能中的不可能性定理由机制基础划分为五种类别——演绎、不可区分性、归纳、权衡与计算困难性,突出了人工智能安全、对齐与可控性方面的根本限制。文章引入了若干新成果,如可解释性的不公平性以及智能体自我意识的约束,强调演绎不可能性排除了100%的安全保障。

ABSTRACT

An impossibility theorem demonstrates that a particular problem or set of problems cannot be solved as described in the claim. Such theorems put limits on what is possible to do concerning artificial intelligence, especially the super-intelligent one. As such, these results serve as guidelines, reminders, and warnings to AI safety, AI policy, and governance researchers. These might enable solutions to some long-standing questions in the form of formalizing theories in the framework of constraint satisfaction without committing to one option. We strongly believe this to be the most prudent approach to long-term AI safety initiatives. In this paper, we have categorized impossibility theorems applicable to AI into five mechanism-based categories: deduction, indistinguishability, induction, tradeoffs, and intractability. We found that certain theorems are too specific or have implicit assumptions that limit application. Also, we added new results (theorems) such as the unfairness of explainability, the first explainability-related result in the induction category. The remaining results deal with misalignment between the clones and put a limit to the self-awareness of agents. We concluded that deductive impossibilities deny 100%-guarantees for security. In the end, we give some ideas that hold potential in explainability, controllability, value alignment, ethics, and group decision-making. They can be deepened by further investigation.

研究动机与目标

  • 基于演绎、归纳与计算困难性等底层机制,系统性地对人工智能中的不可能性定理进行分类。
  • 识别并形式化人工智能安全、对齐与可控性中的固有局限,尤其针对超级智能系统。
  • 突出被忽视或新推导出的结果,包括可解释性的不公平性以及智能体自我意识的约束。
  • 通过提供一种约束满足框架,避免过早承诺特定解决方案,为长期人工智能安全举措提供指导。
  • 通过形式化不可能性约束,激发在可解释性、价值对齐、伦理与群体决策制定方面的进一步研究。

提出的方法

  • 作者将现有不可能性定理划分为五种类别:演绎、不可区分性、归纳、权衡与计算困难性。
  • 分析已知定理的假设与适用范围,识别出那些过于狭窄或隐含受限而难以广泛适用的定理。
  • 本文提出新的理论成果,包括在归纳类别中形式化一个关于可解释性的不可能性结果。
  • 探讨智能体克隆与自我复制对自我意识的影响,揭示此类能力的内在限制。
  • 该框架以约束满足为基础,允许在不承诺特定人工智能架构或安全机制的前提下形式化理论。
  • 理论分析得到对人工智能安全、伦理与治理文献中103篇参考文献的全面综述支持。

实验结果

研究问题

  • RQ1不可能性定理对安全且可控的超级智能人工智能发展施加了哪些根本性限制?
  • RQ2演绎不可能性如何阻止人工智能安全与对齐的100%保障?
  • RQ3可解释性概念在何种程度上面临固有局限?这一局限通过归纳类别中的新不可能性结果得以形式化?
  • RQ4克隆与自我复制在多大程度上限制了智能体的自我意识?
  • RQ5如何利用约束满足框架在不预设具体解决方案的前提下形式化人工智能安全理论?

主要发现

  • 演绎不可能性从根本上阻止了人工智能系统实现100%保证的安全性,确立了可验证安全性的根本上限。
  • 本文提出一个新颖的不可能性结果,证明了在人工智能中可解释性本质上存在不公平性,尤其是在归纳推理的语境下。
  • 当智能体被克隆时,自我意识受到约束,表明此类系统无法实现完全的自我意识。
  • 许多现有不可能性定理过于具体,或依赖隐含假设,限制了其在人工智能安全领域的普遍适用性。
  • 约束满足框架为形式化人工智能安全理论提供了可行路径,而无需预设单一架构或对齐方法。
  • 本综述在可解释性、可控性、价值对齐、伦理与群体决策制定方面识别出有前景的研究方向,所有方向均基于形式化的不可能性结果。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。