[Paper Review] Paradoxes of rational agency and formal systems that verify their own soundness
This paper resolves paradoxes in rational agency by extending Peano arithmetic with an assertibility predicate, enabling formal systems to effectively verify their own soundness. It demonstrates that such systems can overcome limitations in self-trust and delegation, allowing agents to act based on provability without requiring full proof execution, thus resolving issues in AI reasoning and hierarchical agent coordination.
We consider extensions of Peano arithmetic which include an assertibility predicate. Any such system which is arithmetically sound effectively verifies its own soundness. This leads to the resolution of a range of paradoxes involving rational agents who are licensed to act under precisely defined conditions.
Motivation & Objective
- To resolve paradoxes in rational agency where agents cannot trust their own or others' proofs despite knowing they are sound.
- To address the limitations of Löb's theorem and Gödel's second incompleteness theorem in AI systems that must act based on provable conditions.
- To design formal systems that can effectively verify their own soundness, overcoming the inability to affirm the soundness scheme (⁎).
- To enable hierarchical agent systems to delegate actions based on provable outcomes, even when proofs are too long to execute.
Proposed method
- Extends Peano arithmetic with an assertibility predicate to create a system S* that can reason about its own soundness.
- Introduces a hierarchy of agents M₁, M₂, ..., each with distinct licensing criteria based on iterated assertibility operators (□^κᵢ).
- Uses the box rule and provability logic to derive that if M₁ can prove M₂ will act under its criteria, then M₁ can justify activating M₂.
- Employs a formal system S* with an infinite sequence of constants (κᵢ) satisfying κᵢ = κᵢ₊₁ + 1 to model increasingly strong levels of assertibility.
- Applies Theorems 3.1 and 3.2 to show that if a system proves a statement is provable, it can also assert its truth under the assertibility predicate.
- Leverages the fact that the κᵢ values are unspecified large numbers, making the licensing criteria effectively equivalent across agents despite formal differences.
Experimental results
Research questions
- RQ1Can a formal system be constructed such that if it proves a sentence is provable, it can also conclude the sentence is true, thereby verifying its own soundness?
- RQ2How can an AI agent be designed to act based on the provability of a condition without needing to execute the full proof?
- RQ3Can a rational agent trust another agent’s proof even if it cannot verify the proof directly, provided the system is self-verifying?
- RQ4Is it possible to construct a hierarchy of agents where each can delegate actions to the next based on provable success conditions?
- RQ5How can the paradoxes of naturalistic and reflective trust be resolved in formal systems that cannot directly assert their own soundness?
Key findings
- The extended system S* with an assertibility predicate allows agents to effectively verify their own soundness, bypassing Löb’s theorem limitations.
- Agents can be licensed to act based on the provability of a condition, even when the proof is too long to execute, by using the assertibility operator.
- The system enables hierarchical delegation: M₁ can activate M₂ based on M₂’s provable ability to achieve a goal, even if M₁ cannot verify the proof directly.
- The use of non-standard constants (κᵢ) with κᵢ = κᵢ₊₁ + 1 allows for a consistent hierarchy of assertibility levels without violating consistency.
- The system resolves the reflective coherence paradox by ensuring that if a system proves all instances of A(n) are provable, it can infer (∀n)A(n) under the assertibility predicate.
- The approach provides a non-standard but consistent solution to the delegated action problem in AI, where agents must act based on trust in provable outcomes.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.