Skip to main content
QUICK REVIEW

[论文解读] Modeling Moral Choices in Social Dilemmas with Multi-Agent Reinforcement Learning

Elizaveta Tennant, Stephen Hailes|arXiv (Cornell University)|Jan 20, 2023
Evolutionary Game Theory and Cooperation被引用 4
一句话总结

本文提出了一种新颖的框架,通过基于功利主义、义务论和美德伦理的简化道德奖励函数,对多智能体强化学习中的道德决策进行建模。结果表明,基于规范和友善导向的道德智能体能够学会合作,但容易被利用;而以平等为核心的智能体在训练过程中表现出掠夺性行为,混合美德智能体则保持合作但易受利用。

ABSTRACT

Practical uses of Artificial Intelligence (AI) in the real world have demonstrated the importance of embedding moral choices into intelligent agents. They have also highlighted that defining top-down ethical constraints on AI according to any one type of morality is extremely challenging and can pose risks. A bottom-up learning approach may be more appropriate for studying and developing ethical behavior in AI agents. In particular, we believe that an interesting and insightful starting point is the analysis of emergent behavior of Reinforcement Learning (RL) agents that act according to a predefined set of moral rewards in social dilemmas. In this work, we present a systematic analysis of the choices made by intrinsically-motivated RL agents whose rewards are based on moral theories. We aim to design reward structures that are simplified yet representative of a set of key ethical systems. Therefore, we first define moral reward functions that distinguish between consequence- and norm-based agents, between morality based on societal norms or internal virtues, and between single- and mixed-virtue (e.g., multi-objective) methodologies. Then, we evaluate our approach by modeling repeated dyadic interactions between learning moral agents in three iterated social dilemma games (Prisoner's Dilemma, Volunteer's Dilemma and Stag Hunt). We analyze the impact of different types of morality on the emergence of cooperation, defection or exploitation, and the corresponding social outcomes. Finally, we discuss the implications of these findings for the development of moral agents in artificial and mixed human-AI societies.

研究动机与目标

  • 探究功利主义、义务论和美德伦理等不同道德框架如何塑造强化学习智能体在社会困境中的行为。
  • 开发一种系统化的方法,以简化但具代表性的方式编码反映关键伦理体系的道德奖励函数。
  • 分析在三种重复社会困境游戏中,双智能体互动下出现的合作、背叛和被利用行为。
  • 探讨人工智能智能体之间道德多样性对设计伦理型人工社会及人机混合社会的启示。
  • 为未来在复杂多智能体环境中研究道德智能体交互行为提供基础的方法论平台。

提出的方法

  • 基于三种伦理框架设计道德奖励函数:功利主义(最大化全局奖励)、义务论(遵守规范)和美德伦理(强调内在美德,如友善与平等)。
  • 实施内在动机驱动的强化学习智能体,其策略由这些道德奖励函数引导,而非纯粹自利回报。
  • 在三种重复社会困境游戏中评估智能体:囚徒困境、志愿者困境和猎鹿游戏,采用重复的双智能体互动形式。
  • 采用固定对手类型的多智能体强化学习设置,通过100次独立训练运行分析学习动态与社会结果。
  • 引入一种混合美德奖励函数,结合友善与平等并可调节权重,以研究多目标道德行为。
  • 通过分析收敛速度、合作率和被利用模式,比较不同道德智能体类型的表现。

实验结果

研究问题

  • RQ1不同的道德奖励结构(功利主义、义务论、基于美德的)如何影响多智能体强化学习中合作或背叛的出现?
  • RQ2当道德智能体在重复社会困境中互动时,其学习动态与社会结果如何?
  • RQ3在合作与被利用风险方面,混合美德道德智能体(平衡友善与平等)与纯美德智能体相比有何差异?
  • RQ4基于规范的道德智能体(如义务论、美德-友善型)在多大程度上能学会合作策略,但仍然容易被利用?
  • RQ5在多目标道德奖励函数中,不同美德的权重如何影响智能体行为与社会结果?

主要发现

  • 功利主义、义务论以及美德-友善智能体在所有三款游戏中均持续学习到合作策略,实现了帕累托最优的合作结果。
  • 这些基于规范和友善导向的智能体系统性地被自私对手所利用,表明合作并不保证对利用行为具有鲁棒性。
  • 美德-平等智能体在收敛前的训练阶段表现出掠夺性行为,表明以平等为核心的道德体系可能导致合作不够稳定。
  • 混合美德智能体(友善与平等权重相等)的行为与纯友善智能体相似——合作但易受利用,表明在缺乏强平等激励时,友善占主导地位。
  • 仅当平等权重(β)非常高时,混合美德智能体才转向更具防御性且合作性更低的行为,表明道德权衡显著影响策略的形成。
  • 本研究为在人工社会中建模多样化的道德智能体奠定了方法论基础,支持未来在复杂多智能体环境中对道德动态的研究。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。