Skip to main content
QUICK REVIEW

[Paper Review] Impossibility Results in AI: A Survey

Mario Brčić, Roman V. Yampolskiy|arXiv (Cornell University)|Sep 1, 2021
Ethics and Social Impacts of AI4 citations
TL;DR

This survey categorizes impossibility theorems in artificial intelligence into five mechanism-based classes—deduction, indistinguishability, induction, tradeoffs, and intractability—highlighting fundamental limits on AI safety, alignment, and controllability. It introduces novel results such as the unfairness of explainability and constraints on self-awareness in agents, emphasizing that deductive impossibilities preclude 100% security guarantees.

ABSTRACT

An impossibility theorem demonstrates that a particular problem or set of problems cannot be solved as described in the claim. Such theorems put limits on what is possible to do concerning artificial intelligence, especially the super-intelligent one. As such, these results serve as guidelines, reminders, and warnings to AI safety, AI policy, and governance researchers. These might enable solutions to some long-standing questions in the form of formalizing theories in the framework of constraint satisfaction without committing to one option. We strongly believe this to be the most prudent approach to long-term AI safety initiatives. In this paper, we have categorized impossibility theorems applicable to AI into five mechanism-based categories: deduction, indistinguishability, induction, tradeoffs, and intractability. We found that certain theorems are too specific or have implicit assumptions that limit application. Also, we added new results (theorems) such as the unfairness of explainability, the first explainability-related result in the induction category. The remaining results deal with misalignment between the clones and put a limit to the self-awareness of agents. We concluded that deductive impossibilities deny 100%-guarantees for security. In the end, we give some ideas that hold potential in explainability, controllability, value alignment, ethics, and group decision-making. They can be deepened by further investigation.

Motivation & Objective

  • To systematically categorize impossibility theorems in AI based on underlying mechanisms such as deduction, induction, and intractability.
  • To identify and formalize inherent limitations in AI safety, alignment, and controllability, especially for superintelligent systems.
  • To highlight underappreciated or newly derived results, including the unfairness of explainability and constraints on agent self-awareness.
  • To guide long-term AI safety initiatives by providing a constraint-satisfaction framework that avoids premature commitment to specific solutions.
  • To stimulate further research in explainability, value alignment, ethics, and group decision-making under formalized impossibility constraints.

Proposed method

  • The authors classify existing impossibility theorems into five mechanism-based categories: deduction, indistinguishability, induction, tradeoffs, and intractability.
  • They analyze the assumptions and scope of known theorems, identifying those that are too narrow or implicitly constrained for broad application.
  • The paper introduces new theoretical results, including a formal impossibility result on explainability in the induction category.
  • It examines the implications of agent cloning and self-replication for self-awareness, showing inherent limits on such capabilities.
  • The framework is grounded in constraint satisfaction, allowing formalization of theories without committing to specific AI architectures or safety mechanisms.
  • Theoretical analysis is supported by a comprehensive review of 103 references from AI safety, ethics, and governance literature.

Experimental results

Research questions

  • RQ1What are the fundamental limits imposed by impossibility theorems on the development of safe and controllable superintelligent AI?
  • RQ2How do deductive impossibilities prevent 100% guarantees in AI security and alignment?
  • RQ3In what ways does the concept of explainability face inherent limitations, as formalized by a new impossibility result in the induction category?
  • RQ4To what extent do cloning and self-replication constrain the self-awareness of intelligent agents?
  • RQ5How can constraint satisfaction frameworks be used to formalize AI safety theories without predefining specific solutions?

Key findings

  • Deductive impossibilities inherently prevent 100%-guaranteed security in AI systems, establishing a fundamental upper bound on verifiable safety.
  • The paper introduces a novel impossibility result demonstrating the inherent unfairness of explainability in AI, particularly in the context of inductive reasoning.
  • Constraints on self-awareness arise when agents are cloned, showing that full self-awareness cannot be achieved in such systems.
  • Many existing impossibility theorems are too specific or rely on implicit assumptions that limit their general applicability to AI safety.
  • The framework of constraint satisfaction offers a viable path for formalizing AI safety theories without committing to a single architectural or alignment approach.
  • The survey identifies promising research directions in explainability, controllability, value alignment, ethics, and group decision-making, all grounded in formal impossibility results.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.