[Paper Review] Using experimental game theory to transit human values to ethical AI
This paper proposes a computational framework that uses experimental game theory to extract and embed human moral values—particularly Kantian ethics—into AI systems, enabling ethical decision-making through quantifiable human behavior data. By modeling AI-human interactions in strategy and outcome spaces, it demonstrates how to test AI ethics via iterated Prisoner’s Dilemma experiments, showing that algorithms like Generosity are ethically sound while Extortion is not, offering a rigorous, testable, and industry-conductible method for ethical AI design.
Knowing the reflection of game theory and ethics, we develop a mathematical representation to bridge the gap between the concepts in moral philosophy (e.g., Kantian and Utilitarian) and AI ethics industry technology standard (e.g., IEEE P7000 standard series for Ethical AI). As an application, we demonstrate how human value can be obtained from the experimental game theory (e.g., trust game experiment) so as to build an ethical AI. Moreover, an approach to test the ethics (rightness or wrongness) of a given AI algorithm by using an iterated Prisoner's Dilemma Game experiment is discussed as an example. Compared with existing mathematical frameworks and testing method on AI ethics technology, the advantages of the proposed approach are analyzed.
Motivation & Objective
- To develop a mathematically rigorous, computable, and experimentally testable method for aligning AI behavior with human values, especially Kantian ethics.
- To bridge the gap between moral philosophy (Kantian and Utilitarian) and AI ethics industry standards (e.g., IEEE P7000) through a unified mathematical representation.
- To enable the quantitative transfer of human values—such as fairness, trust, and altruism—from controlled behavioral experiments into ethical AI design.
- To provide a practical, scalable, and testable method for evaluating the ethicality of AI algorithms using outcome space analysis in non-cooperative games.
- To overcome limitations of inductive, linguistically based AI ethics standards by introducing a deductive, data-driven, and mathematically structured approach.
Proposed method
- The paper introduces a mathematical representation of AI-human interaction: $\mathbf{S}^{a}_{i} \otimes \mathbf{S}^{b}_{j} \rightarrow \mathbf{O}$, where $\mathbf{S}^{a}$ and $\mathbf{S}^{b}$ are strategy spaces for AI and human, and $\mathbf{O}$ is the outcome space representing payoffs (e.g., rewards, costs, fairness).
- It defines Kantian value as a non-utilitarian, ethically grounded component of behavior that deviates from pure rational choice, and embeds it into the outcome space as a constraint or target for AI alignment.
- Human values such as trust and fairness are extracted from experimental game theory data (e.g., trust game, public goods game), which are then used as parameters to calibrate ethical AI behavior.
- The ethicality of an AI algorithm is tested using the iterated Prisoner’s Dilemma (IPD), where the outcome space is analyzed to determine if the AI’s behavior is fair and non-exploitative.
- The method uses theoretical outcome lines (e.g., red line for Extortion, green for Generosity) to compare actual experimental scores against ethical benchmarks, with fairness and efficiency as key criteria.
- It contrasts its paradigm with prior models by explicitly defining Kantian value in the outcome space, enabling computable solutions, unlike previous models where ethical decisions remained ambiguous.
Experimental results
Research questions
- RQ1How can Kantian and Utilitarian moral philosophies be formally integrated into a single computational framework for AI ethics?
- RQ2Can human values derived from experimental game theory be reliably used to calibrate ethical AI behavior in non-cooperative settings?
- RQ3How can the ethicality of a given AI algorithm be quantitatively assessed using outcome space analysis in repeated games?
- RQ4What are the limitations of current inductive, linguistically based AI ethics standards (e.g., IEEE P7000), and how can they be improved with a deductive, data-driven approach?
- RQ5To what extent can experimental data from trust and public goods games be used to define ethical control parameters for AI agents?
Key findings
- The Generosity strategy in the iterated Prisoner’s Dilemma, which satisfies $\frac{3 - s_{hg}}{3 - s_g} = \frac{1}{3}$, achieves maximum scores for both AI and human (3), indicating fairness and efficiency, and is therefore ethically sound.
- The Extortion strategy, which satisfies $\frac{s_{he} - 1}{s_e - 1} = \frac{1}{3}$, results in a maximum human score of 1.907 and AI score of 3.727, creating an unfair outcome and thus being deemed unethical.
- Experimental data from human subjects in non-cooperative games (e.g., trust game) can be used to extract quantifiable values of fairness, trust, and altruism, which can then be embedded into AI decision-making.
- The proposed framework enables a computable, testable, and mathematically rigorous method to evaluate AI ethics, overcoming the ambiguity and subjectivity of linguistic checklists used in standards like IEEE P7000.
- By explicitly modeling Kantian values in the outcome space, the method resolves the conflict between Kantian and Utilitarian ethics in AI, allowing for ethical behavior that goes beyond pure utility maximization.
- The approach is scalable and applicable to real-world AI systems, including autonomous vehicles, drones, and financial agents, by using behavioral data to guide ethical design.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.