Skip to main content
QUICK REVIEW

[论文解读] The emergence of division of labor through decentralized social sanctioning

Anil Yaman, Joel Z. Leibo|arXiv (Cornell University)|Aug 10, 2022
Evolutionary Game Theory and Cooperation参考文献 59被引用 4
一句话总结

本文提出,去中心化的社会惩罚机制——以涌现的社会规范形式建模——可使自利的、终身学习的个体自主发展出高效的分工体系,即使某些角色的内在回报较低。通过采用基于规范的奖励调整的epsilon-greedy强化学习框架,该模型在无集中协调的情况下实现了稳定且高群体福利的角色分配,表明社会规范可通过去中心化的激励对齐解决集体行动困境。

ABSTRACT

Human ecological success relies on our characteristic ability to flexibly self-organize into cooperative social groups, the most successful of which employ substantial specialization and division of labor. Unlike most other animals, humans learn by trial and error during their lives what role to take on. However, when some critical roles are more attractive than others, and individuals are self-interested, then there is a social dilemma: each individual would prefer others take on the critical but unremunerative roles so they may remain free to take one that pays better. But disaster occurs if all act thusly and a critical role goes unfilled. In such situations learning an optimum role distribution may not be possible. Consequently, a fundamental question is: how can division of labor emerge in groups of self-interested lifetime-learning individuals? Here we show that by introducing a model of social norms, which we regard as emergent patterns of decentralized social sanctioning, it becomes possible for groups of self-interested individuals to learn a productive division of labor involving all critical roles. Such social norms work by redistributing rewards within the population to disincentivize antisocial roles while incentivizing prosocial roles that do not intrinsically pay as well as others.

研究动机与目标

  • 解决自利个体如何在群体中学习承担必要但回报较低的角色这一挑战。
  • 探究去中心化的社会规范——即涌现的惩罚模式——是否能激励亲社会的角色选择。
  • 建模社会规范如何重新分配奖励,以抑制反社会角色并促进关键但薪酬较低的角色。
  • 比较基于规范的学习与自私、全局及利他主义学习策略在实现群体层面分工方面的有效性。

提出的方法

  • 个体通过epsilon-greedy Q-learning算法(第一阶段)学习角色,基于经验估计的收益选择角色。
  • 在第二阶段,社会规范被建模为k×k的鼓励/抑制规则矩阵,根据观察到的行为动态调整奖励。
  • 社会惩罚以去中心化方式实施:拥有足够奖励的个体可执行规范,而资源匮乏的个体则不能,从而保持规范记忆。
  • 规范演化过程采用爬山算法,并引入稀疏性惩罚项(λ||S||₀),以优化社会规范以实现最大群体收益。
  • 模型通过两种游戏评估性能:定居点维护与公共牧场,收益由集体角色分配决定。
  • 对比基线包括自私(epsilon-greedy)、利他(加权邻近群体收益)以及全局优化(基于空间角色分布的遗传算法)策略。

实验结果

研究问题

  • RQ1去中心化的社会惩罚机制能否使自利个体学习到包含所有关键角色的分工体系?
  • RQ2在缺乏内在回报的情况下,重新分配奖励的规范规则如何影响亲社会角色选择的出现?
  • RQ3所提出的基于规范的学习方法是否优于自私、利他或全局优化策略,在实现高群体福利方面表现更优?
  • RQ4资源约束如何影响社会规范在角色分配中的实施与持续性?

主要发现

  • 基于规范的学习方法在定居点维护与公共牧场两种游戏中,均成功实现了稳定且高群体福利的角色分配。
  • 模型表明,社会规范能有效抑制搭便车行为,并鼓励参与回报较低但至关重要的角色。
  • 去中心化惩罚优于自私学习策略,后者因收益不均而未能填补关键角色。
  • 在规范演化中引入稀疏性惩罚(λ||S||₀)促进了更简单、更具可解释性的规范体系,同时仍保持高绩效。
  • 利他主义及自私/利他混合策略表现居中,但未能持续实现最优角色分配,不如基于规范的方法。
  • 即使在某些角色缺失时,模型仍能保留规范性规则,表明规范具有独立于当前角色可用性的文化记忆。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。