Skip to main content
QUICK REVIEW

[论文解读] Modeling the Formation of Social Conventions in Multi-Agent Populations.

Ismael T. Freire, Clément Moulin-Frier|arXiv (Cornell University)|Feb 16, 2018
Experimental Behavioral Economics Studies参考文献 51被引用 8
一句话总结

本文在分布式自适应控制(DAC)理论框架内提出了一种基于控制的强化学习(CRL)架构,将反应式感知运动控制与无模型强化学习相结合,使多智能体系统能够学习并形成社会规范。CRL框架在博弈论任务中成功实现了最优协调,复现了人类实验数据在奖励效率、公平性及规范稳定性方面的表现,涵盖离散时间与连续时间两种情形。

ABSTRACT

In order to understand the formation of social conventions we need to know the specific role of control and learning in multi-agent systems. To advance in this direction, we propose, within the framework of the Distributed Adaptive Control (DAC) theory, a novel Control-based Reinforcement Learning architecture (CRL) that can account for the acquisition of social conventions in multi-agent populations that are solving a benchmark social decision-making problem. Our new CRL architecture, as a concrete realization of DAC multi-agent theory, implements a low-level sensorimotor control loop handling the agent's reactive behaviors (pre-wired reflexes), along with a layer based on model-free reinforcement learning that maximizes long-term reward. We apply CRL in a multi-agent game-theoretic task in which coordination must be achieved in order to find an optimal solution. We show that our CRL architecture is able to both find optimal solutions in discrete and continuous time and reproduce human experimental data on standard game-theoretic metrics such as efficiency in acquiring rewards, fairness in reward distribution and stability of convention formation.

研究动机与目标

  • 理解控制与学习机制如何驱动多智能体系统中的社会规范形成。
  • 填补建模智能体如何在动态环境中通过协调决策获取规范的空白。
  • 开发DAC理论的切实实现,支持反应式行为与长期奖励最大化。
  • 在标准博弈论任务中,以人类行为数据验证该模型。
  • 展示规范形成在离散时间与连续时间设定下的鲁棒性。

提出的方法

  • CRL架构实施了分层控制结构,其中低层为感知运动控制回路,用于实现反应式行为。
  • 无模型强化学习层并行运行,以最大化长期累积奖励。
  • 反应式控制与学习的融合使系统能够对环境反馈做出自适应响应。
  • 该框架被应用于基准多智能体博弈论任务,要求通过协调实现最优结果。
  • 采用标准博弈论指标(如奖励效率、公平性及规范稳定性)对系统进行评估。
  • 在离散时间与连续时间领域对架构进行测试,以评估其泛化能力。

实验结果

研究问题

  • RQ1控制理论框架如何使多智能体系统在协调任务中学习并形成社会规范?
  • RQ2CRL架构在多大程度上复现了人类在奖励效率与公平性方面的实验数据?
  • RQ3CRL智能体在重复试验中形成的规范在稳定性和一致性方面如何?
  • RQ4CRL框架是否能在离散时间与连续时间设定下均实现最优解?
  • RQ5反应式控制与强化学习的融合在规范形成过程中起到何种作用?

主要发现

  • CRL架构在离散时间与连续时间设定下均成功实现了最优解。
  • 该系统复现了人类在获取奖励效率方面的实验数据,表明行为高度一致。
  • 该框架在奖励分配方面表现出公平性,与人类在社会协调任务中的倾向一致。
  • CRL系统中的规范形成表现出高度的时间稳定性,与人类数据一致。
  • 感知运动控制与强化学习的融合使规范获取更加稳健且具备自适应能力。
  • CRL模型在时间域间具有泛化能力,在不同时间设定下均表现出一致性能。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。