Skip to main content
QUICK REVIEW

[论文解读] Impossibility of deducing preferences and rationality from human policy.

Stuart Armstrong, Sören Mindermann|arXiv (Cornell University)|Dec 15, 2017
Game Theory and Applications参考文献 31被引用 9
一句话总结

本文证明,在缺乏关于人类理性的规范性假设的前提下,仅凭行为无法推导出人类的奖励函数,即使在理想数据条件下也是如此。本文建立了一个自由定理,表明从行为来看,奖励函数与规划算法组件在本质上是不可区分的,这使得标准逆强化学习在缺乏额外规范性约束的情况下不可行。

ABSTRACT

Inverse reinforcement learning (IRL) attempts to infer human rewards or preferences from observed behavior. However, human planning systematically deviates from rationality. Though there has been some IRL work which assumes humans are noisily rational, there has been little analysis of the general problem of inferring the reward of a human of unknown rationality. The observed behavior can, in principle, be decomposed into two composed into two components: a reward function and a planning algorithm that maps reward function to policy. Both of these variables have to be inferred from behaviour. This paper presents a Free theorem in this area, showing that, without making `normative' assumptions beyond the data, nothing about the human reward function can be deduced from human behaviour. Unlike most No Free Lunch theorems, this cannot be alleviated by regularising with simplicity assumptions. The simplest hypotheses are generally degenerate. The paper will then sketch how one might begin to use normative assumptions to get around the problem, without which solving the general IRL problem is impossible. The reward function-planning algorithm formalism can also be used to encode what it means for an agent to manipulate or override human preferences.

研究动机与目标

  • 研究在逆强化学习中,从观察到的行为推断人类偏好的理论极限。
  • 识别当人类规划偏离理性时,标准IRL方法为何会失效。
  • 形式化行为分解为奖励函数与规划算法的组合,作为偏好推断的核心挑战。
  • 证明在仅观察到行为的基础上,若无规范性假设,则无法对人类奖励做出任何推断。
  • 探讨规范性假设在实践中如何克服这一理论上的不可能性。

提出的方法

  • 将人类行为形式化为奖励函数与规划算法的组合,二者均映射到一个策略。
  • 提出一个自由定理,证明在最小假设下,仅从行为中无法推导出任何关于奖励函数的信息。
  • 分析简化性假设在正则化中的作用,并表明其无法解决由于退化假设导致的不确定性。
  • 利用奖励-规划形式化来建模代理对人类偏好的操纵或覆盖。
  • 证明该问题无法通过数据或简化性本身解决,必须引入规范性约束才能推进。

实验结果

研究问题

  • RQ1在无额外假设的前提下,能否唯一地从观察到的行为中推断出人类奖励函数?
  • RQ2当人类规划非理性时,逆强化学习的根本理论限制是什么?
  • RQ3通过简化性假设进行正则化在多大程度上能解决奖励与规划组件之间的模糊性?
  • RQ4如何有意义地引入规范性假设,以使IRL在面对这一不可能性时仍可行?
  • RQ5奖励-规划形式化在何种方式下可用来建模代理对人类偏好的操纵?

主要发现

  • 在缺乏关于理性的规范性假设的前提下,仅从行为中推导人类奖励函数在理论上是不可能的。
  • 自由定理表明,多种奖励与规划算法的组合可产生完全相同的行为,导致推断问题不可判定。
  • 基于简化的正则化方法失效,因为最简化的假设通常为退化且无法识别的。
  • 奖励-规划形式化提供了一个建模偏好操纵的框架,其中代理可覆盖或扭曲人类偏好。
  • 若无规范性假设,解决一般性的逆强化学习问题在根本上是不可能的。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。