Skip to main content
QUICK REVIEW

[论文解读] How and Why to Manipulate Your Own Agent: On the Incentives of Users of Learning Agents

Yoav Kolumbus, Noam Nisan|arXiv (Cornell University)|Dec 14, 2021
Auction Theory and Applications被引用 5
一句话总结

本文提出了一种元博弈框架,用户通过谎报其学习智能体的私有参数,战略性地操纵智能体以在重复博弈中提升长期效用。研究表明,用户可通过此类操纵获益,尤其在使用后悔最小化智能体(如乘法权重或跟随扰动领导者)时,导致的结果与诚实报告和标准纳什均衡显著偏离。

ABSTRACT

The usage of automated learning agents is becoming increasingly prevalent in many online economic applications such as online auctions and automated trading. Motivated by such applications, this paper is dedicated to fundamental modeling and analysis of the strategic situations that the users of automated learning agents are facing. We consider strategic settings where several users engage in a repeated online interaction, assisted by regret-minimizing learning agents that repeatedly play a "game" on their behalf. We propose to view the outcomes of the agents' dynamics as inducing a "meta-game" between the users. Our main focus is on whether users can benefit in this meta-game from "manipulating" their own agents by misreporting their parameters to them. We define a general framework to model and analyze these strategic interactions between users of learning agents for general games and analyze the equilibria induced between the users in three classes of games. We show that, generally, users have incentives to misreport their parameters to their own agents, and that such strategic user behavior can lead to very different outcomes than those anticipated by standard analysis.

研究动机与目标

  • 建模用户在重复博弈中委托学习智能体决策时的战略激励。
  • 分析用户通过谎报其私有参数(如效用)来操纵自身智能体是否能获益。
  • 研究后悔最小化智能体(如乘法权重、FTPL)的动态如何在用户之间诱导出元博弈,其中用户行动为参数声明,结果为长期平均效用。
  • 刻画该元博弈中的均衡,并与标准博弈论解(如纳什均衡)进行比较。
  • 评估不同学习算法(如乘法权重、FTPL)对用户激励和结果的影响。

提出的方法

  • 形式化一个元博弈,其中用户的策略是向其智能体声明参数,收益为智能体动态产生的长期平均效用。
  • 使用后悔最小化算法(如乘法权重(MW)和跟随扰动领导者(FTPL))建模智能体行为,这些算法收敛至广义相关均衡(CCE)。
  • 通过将智能体的经验行为视为用户报告参数的函数,分析由此产生的用户均衡。
  • 通过三类博弈(包括对立利益博弈)的模拟,比较诚实报告与操纵报告下的结果。
  • 通过时间平均策略频率估计智能体动态向纳什均衡的收敛性,并测量误差为 O(1/√T)(针对MW智能体)。
  • 将模拟结果与理论预测进行比较,包括与纳什均衡的偏离及操纵带来的效用增益。

实验结果

研究问题

  • RQ1用户能否战略性地向其学习智能体谎报其私有参数(如效用),以获得更高的长期收益?
  • RQ2后悔最小化智能体(如MW、FTPL)的动态如何影响用户之间诱导出的元博弈中的均衡结果?
  • RQ3操纵性声明在多大程度上导致与基础博弈纳什均衡不同的结果?
  • RQ4学习算法的更新规则(如步长 η)在塑造用户激励和向均衡收敛方面起什么作用?
  • RQ5是否存在操纵配置,使得所有用户相比诚实报告均实现帕累托改进?

主要发现

  • 用户有强烈动机向其学习智能体谎报参数,因为此类操纵可带来高于诚实报告的长期效用。
  • 在对立利益博弈示例中,操纵配置(c=d=1)使双方用户收益均高于诚实声明,尽管该配置并非元博弈均衡。
  • 模拟结果表明,乘法权重和FTPL智能体在操纵下产生相似的效用结果,且平均效用接近各操纵配置的理论预测值。
  • 对于乘法权重智能体,时间平均策略频率与纳什均衡分布的偏差随 O(1/√T) 减小,证实了理论收敛速率。
  • 1,000次模拟运行的策略频率直方图在纳什均衡附近高度集中,表明在诚实报告下长期行为稳定。
  • 操纵下智能体的经验动态导致结果与标准博弈论均衡显著不同,凸显了在智能体层面分析之外建模用户激励的重要性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。