Skip to main content
QUICK REVIEW

[论文解读] Modeling Theory of Mind in Multi-Agent Games Using Adaptive Feedback Control

Ismael T. Freire, Xerxes D. Arsiwalla|arXiv (Cornell University)|May 29, 2019
Evolutionary Game Theory and Cooperation参考文献 37被引用 9
一句话总结

本文提出了一种基于控制理论的框架,用于在多智能体系统中建模心智理论(ToM),采用自适应反馈控制,整合分层控制层级中的自上而下预测与自下而上误差信号。结果表明,在博弈论任务中,建模他人策略或采取理性行为的智能体优于纯强化学习智能体,表现出更快的收敛速度和更高的性能,尤其在协调与竞争类博弈(如猎鹿博弈和情人对决)中表现突出。

ABSTRACT

A major challenge in cognitive science and AI has been to understand how autonomous agents might acquire and predict behavioral and mental states of other agents in the course of complex social interactions. How does such an agent model the goals, beliefs, and actions of other agents it interacts with? What are the computational principles to model a Theory of Mind (ToM)? Deep learning approaches to address these questions fall short of a better understanding of the problem. In part, this is due to the black-box nature of deep networks, wherein computational mechanisms of ToM are not readily revealed. Here, we consider alternative hypotheses seeking to model how the brain might realize a ToM. In particular, we propose embodied and situated agent models based on distributed adaptive control theory to predict actions of other agents in five different game theoretic tasks (Harmony Game, Hawk-Dove, Stag-Hunt, Prisoner's Dilemma and Battle of the Exes). Our multi-layer control models implement top-down predictions from adaptive to reactive layers of control and bottom-up error feedback from reactive to adaptive layers. We test cooperative and competitive strategies among seven different agent models (cooperative, greedy, tit-for-tat, reinforcement-based, rational, predictive and other's-model agents). We show that, compared to pure reinforcement-based strategies, probabilistic learning agents modeled on rational, predictive and other's-model phenotypes perform better in game-theoretic metrics across tasks. Our autonomous multi-agent models capture systems-level processes underlying a ToM and highlight architectural principles of ToM from a control-theoretic perspective.

研究动机与目标

  • 理解自主智能体中心智理论(ToM)的架构与计算原理。
  • 开发一种生物上合理、具身化且情境化的认知架构,以支持多智能体交互中的社会推理。
  • 检验基于控制的模型是否能通过预测与自适应反馈,在社会决策任务中优于标准强化学习。
  • 在五种经典博弈论场景中,于仿真环境与真实具身机器人环境中验证该模型。
  • 探索在共享社会动态下,不同行为表型(如理性型、预测型及其他智能体建模型)如何涌现并表现。

提出的方法

  • 设计一种多层控制架构,包含通过反馈回路连接的反应式(自下而上)与自适应(自上而下)控制层。
  • 在自适应层实施自上而下的预测,以指导反应式行为,同时利用自下而上的误差信号更新自适应策略。
  • 在自适应层使用强化学习以最大化长期奖励,其依据来自反应式层的预测误差。
  • 采用博弈论基准测试(和谐博弈、猎鹿博弈、鹰鸽博弈、囚徒困境、情人对决)评估智能体性能。
  • 在仿真环境与真实机器人平台(使用ePuck机器人,具备部分可观测性与实时传感器反馈)中验证模型。
  • 比较七类智能体:合作型、贪婪型、以牙还牙型、基于强化学习型、理性型、预测型及其他智能体建模型。

实验结果

研究问题

  • RQ1如何利用生物上合理的控制机制在人工智能体中实现心智理论?
  • RQ2自适应反馈控制架构是否能在社会博弈中实现比纯强化学习更快的收敛速度与更优性能?
  • RQ3能够预测或建模他人策略的智能体是否在性能上优于仅依赖自我优化的智能体?
  • RQ4具身化与情境化条件如何影响多智能体系统中类ToM行为的涌现?
  • RQ5预测型与他人建模型智能体在多大程度上复现了人类在社会决策任务中的行为模式?

主要发现

  • 在所有五项博弈论任务中,采用预测型或他人建模型表型的智能体在奖励与收敛速度方面均优于纯强化学习智能体。
  • 理性智能体(在给定对手动作下假设采取最优行动)在多数博弈中实现了最高性能,表明战略预见能力的价值。
  • 预测型与他人建模型智能体成功学习到在所有博弈类型中预测对手动作与策略,展现出功能性心智理论能力。
  • 在具备实时反馈的真实机器人环境中,克服了在弹道式(非交互式)仿真中出现的收敛问题,表明情境化交互的优势。
  • 在‘情人对决’博弈中,人类的惊讶度分布与预测型及他人建模型智能体高度吻合,而与理性型或纯强化学习智能体不一致,表明其认知合理性。
  • 纯强化学习智能体在复杂协调与竞争博弈中无法收敛至最优均衡,凸显了在缺乏他人建模能力时社会推理的局限性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。