Skip to main content
QUICK REVIEW

[论文解读] Learning and Sustaining Shared Normative Systems via Bayesian Rule Induction in Markov Games

Ninell Oldenburg, Tan Zhi‐Xuan|arXiv (Cornell University)|Feb 20, 2024
Bayesian Modeling and Causal Inference被引用 4
一句话总结

该论文提出了一种贝叶斯规则归纳框架,用于马尔可夫博弈中的多智能体强化学习,使智能体能够通过近似推理来学习并维持共享的规范体系,涵盖义务性规范与禁止性规范。通过假设共享规范性,并基于观察到的合规/违规模式更新信念,智能体实现了快速、样本高效的规范学习与稳定的合作,在长时程环境中优于无模型基线方法。

ABSTRACT

A universal feature of human societies is the adoption of systems of rules and norms in the service of cooperative ends. How can we build learning agents that do the same, so that they may flexibly cooperate with the human institutions they are embedded in? We hypothesize that agents can achieve this by assuming there exists a shared set of norms that most others comply with while pursuing their individual desires, even if they do not know the exact content of those norms. By assuming shared norms, a newly introduced agent can infer the norms of an existing population from observations of compliance and violation. Furthermore, groups of agents can converge to a shared set of norms, even if they initially diverge in their beliefs about what the norms are. This in turn enables the stability of the normative system: since agents can bootstrap common knowledge of the norms, this leads the norms to be widely adhered to, enabling new entrants to rapidly learn those norms. We formalize this framework in the context of Markov games and demonstrate its operation in a multi-agent environment via approximately Bayesian rule induction of obligative and prohibitive norms. Using our approach, agents are able to rapidly learn and sustain a variety of cooperative institutions, including resource management norms and compensation for pro-social labor, promoting collective welfare while still allowing agents to act in their own interests.

研究动机与目标

  • 为解决在多智能体环境中使自主智能体学习并遵守去中心化、共享的社会规范的挑战。
  • 将规范学习建模为一种理性的贝叶斯推理过程,从观察行为中推断结构化规则,而非依赖于反应式或无模型学习。
  • 证明智能体可通过共享信念的形成,自举建立对规范的共同知识,从而实现新智能体的快速引入。
  • 将规范体系形式化为规则结构(义务与禁止),并将其整合进马尔可夫博弈中,以实现协调的、目标导向的行为。
  • 在长时程多智能体环境中评估该框架,展示规范合作的样本效率与稳定性。

提出的方法

  • 通过在标准MDP中增加规范约束(包括禁止性与义务性规则),形式化规范增强的马尔可夫博弈。
  • 将智能体建模为根据当前规范信念在奖励最大化与规范满足模式之间切换的规范合规规划。
  • 实现近似贝叶斯推理,利用他人行为的观察结果,更新候选规范的后验概率。
  • 采用概率规则归纳机制,将规范内容视为假设,其似然度基于观察到的合规或违规模式。
  • 将规范信念整合进策略学习中,使智能体能够通过结构化、可解释的规则,平衡自身利益与规范合规。
  • 利用符号化规则表示,实现规范在智能体间的泛化、通信与可解释性。
(a) Norm-Augmented Markov Game
(a) Norm-Augmented Markov Game

实验结果

研究问题

  • RQ1智能体能否通过从多智能体环境中观察数据推断规则结构的贝叶斯推理,学习共享的社会规范?
  • RQ2与无模型方法相比,假设共享规范性在多大程度上能加速规范学习并提高样本效率?
  • RQ3当智能体最初对规则持有不同信念时,它们在多大程度上能收敛到共同的规范体系?
  • RQ4在无集中式强制的长时程多智能体环境中,规范行为能否出现并持续存在?
  • RQ5与基于习惯或反应式学习相比,符号化、基于规则的规范在可解释性、泛化能力与稳定性方面表现如何?

主要发现

  • 采用近似贝叶斯规则归纳的智能体比无模型基线方法显著更快地学习并维持合作规范,如资源管理与对利他劳动的补偿。
  • 该方法的样本效率比无模型规范学习方法高出数个数量级,收敛所需经验远少于基线。
  • 共享规范性促成了规范的共同知识:一旦达到足够多智能体推断出相同规范,合规性即趋于稳定,新智能体可迅速学习。
  • 该框架通过在以奖励为导向与以义务为导向的模式间切换,成功支持了规范合规规划,确保了自身利益与规范合规的兼顾。
  • 符号化规则表示实现了规范的可解释性、泛化能力与可通信性,支持比纯反应式学习更丰富的规范认知。
  • 在DeepMind的Melting Pot模拟器中的实证评估证实了该方法在复杂、长时程多智能体环境中的可行性与鲁棒性。
(b) Social Norms
(b) Social Norms

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。