Skip to main content
QUICK REVIEW

[论文解读] Model-free conventions in multi-agent reinforcement learning with heterogeneous preferences

Raphaël Koster, Kevin R. McKee|arXiv (Cornell University)|Oct 18, 2020
Reinforcement Learning in Robotics参考文献 74被引用 11
一句话总结

该论文表明,在具有异质偏好的多智能体系统中,无模型强化学习可自发产生复杂且稳定的协调惯例——例如在觅食游戏中出现的集体单一种植状态——而无需依赖基于模型的规划或共同知识。尽管个体智能体具有冲突的内在偏好,它们仍通过习惯性学习自发建立并维持这些惯例,即使偏离这些惯例需要付出巨大的偏好补偿代价,也能实现纳什均衡与帕累托最优结果。

ABSTRACT

Game theoretic views of convention generally rest on notions of common knowledge and hyper-rational models of individual behavior. However, decades of work in behavioral economics have questioned the validity of both foundations. Meanwhile, computational neuroscience has contributed a modernized 'dual process' account of decision-making where model-free (MF) reinforcement learning trades off with model-based (MB) reinforcement learning. The former captures habitual and procedural learning while the latter captures choices taken via explicit planning and deduction. Some conventions (e.g. international treaties) are likely supported by cognition that resonates with the game theoretic and MB accounts. However, convention formation may also occur via MF mechanisms like habit learning; though this possibility has been understudied. Here, we demonstrate that complex, large-scale conventions can emerge from MF learning mechanisms. This suggests that some conventions may be supported by habit-like cognition rather than explicit reasoning. We apply MF multi-agent reinforcement learning to a temporo-spatially extended game with incomplete information. In this game, large parts of the state space are reachable only by collective action. However, heterogeneity of tastes makes such coordinated action difficult: multiple equilibria are desirable for all players, but subgroups prefer a particular equilibrium over all others. This creates a coordination problem that can be solved by establishing a convention. We investigate start-up and free rider subproblems as well as the effects of group size, intensity of intrinsic preference, and salience on the emergence dynamics of coordination conventions. Results of our simulations show agents establish and switch between conventions, even working against their own preferred outcome when doing so is necessary for effective coordination.

研究动机与目标

  • 探究无模型强化学习是否能在具有异质偏好的多智能体系统中支持复杂社会惯例的出现。
  • 研究在信息不完全、集体行动需求和个体偏好冲突的环境中,惯例如何形成。
  • 评估此类惯例是否为纳什均衡与帕累托最优,即使智能体具有不同的内在奖励。
  • 分析群体规模、偏好强度和显著性对惯例形成与稳定性的影响。

提出的方法

  • 设计了一个时空扩展的多智能体强化学习环境,其中智能体以不同颜色的浆果为目标进行觅食,且具有异质的内在偏好。
  • 智能体使用无模型深度Q-learning来基于即时奖励学习策略,无需规划或建模环境。
  • 环境包含一种单一种植机制,智能体可重新种植浆果,从而改变主导颜色并影响集体生长速率。
  • 奖励函数的设计使得智能体在消耗其偏好的浆果颜色时获得更高奖励,但重新种植会降低主导颜色的生长速率。
  • 理论分析推导出单一种植状态成为纳什均衡的条件,基于主导浆果产量损失与偏好颜色消费收益之间的权衡。
  • 模拟实验通过改变群体规模、偏好强度(t)和显著性,研究惯例形成的动态过程与稳定性。

实验结果

研究问题

  • RQ1在缺乏基于模型的规划的情况下,无模型多智能体强化学习是否能产生稳定且大规模的协调惯例?
  • RQ2在存在异质内在偏好时,智能体在何种条件下仍会收敛至单一种植状态?
  • RQ3群体规模、内在偏好的强度以及显著性如何影响惯例的形成与稳定性?
  • RQ4即使智能体必须为协调而牺牲个人偏好,所形成的惯例是否仍为纳什均衡与帕累托最优?

主要发现

  • 当偏好浆果的相对奖励不足以补偿主导浆果产量的损失时,所有智能体协调于单一浆果颜色的单一种植状态会作为纳什均衡出现。
  • 理论分析证实,当主导浆果生长率的损失超过偏好颜色消费带来的收益时,单一种植状态即为纳什均衡,即使偏好比例适中(t=2)亦然。
  • 当群体中至少存在两种不同的群体目标时,单一种植状态为帕累托最优,因为不存在其他状态能同时提升所有共享同一目标的智能体的收益。
  • 智能体频繁在不同惯例间切换,甚至可能为维持协调而放弃其偏好结果,表明其对个体偏好的鲁棒性。
  • 群体规模越大、偏好强度越高,惯例的稳定性越强;而显著性则加速向共享惯例的收敛。
  • 无模型方法成功解决了大规模、信息不完全环境中的深层协调问题,无需显式规划或共同知识。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。