[论文解读] Human irrationality: both bad and good for reward inference
该论文表明,当人类非理性行为被正确建模时,可通过提高行为与奖励之间的互信息,增强逆强化学习中的奖励推理性能,甚至优于理性人类行为。通过将非理性行为形式化为马尔可夫决策过程(MDP)中对贝尔曼最优方程的偏离,作者证明了短视或有偏的决策行为可使人类示范比理性行为更具信息量,尤其当机器人使用了对非理性类型准确的模型时更为显著。
Assuming humans are (approximately) rational enables robots to infer reward functions by observing human behavior. But people exhibit a wide array of irrationalities, and our goal with this work is to better understand the effect they can have on reward inference. The challenge with studying this effect is that there are many types of irrationality, with varying degrees of mathematical formalization. We thus operationalize irrationality in the language of MDPs, by altering the Bellman optimality equation, and use this framework to study how these alterations would affect inference. We find that wrongly modeling a systematically irrational human as noisy-rational performs a lot worse than correctly capturing these biases -- so much so that it can be better to skip inference altogether and stick to the prior! More importantly, we show that an irrational human, when correctly modelled, can communicate more information about the reward than a perfectly rational human can. That is, if a robot has the correct model of a human's irrationality, it can make an even stronger inference than it ever could if the human were rational. Irrationality fundamentally helps rather than hinder reward inference, but it needs to be correctly accounted for.
研究动机与目标
- 研究在不同机器人模型下,人类非理性行为对奖励推理性能的影响。
- 确定非理性人类行为是否可比理性行为为奖励学习提供更多信息。
- 评估正确建模非理性人类与错误地将其建模为噪声理性行为之间的性能差距。
- 探索对非理性行为的近似模型是否仍优于标准的噪声理性假设。
- 评估当机器人正确建模人类非理性行为时,奖励推理性能的理论上限。
提出的方法
- 将非理性行为形式化为马尔可夫决策过程(MDP)中对贝尔曼最优方程的偏离,例如不完全最大化或有偏的转移函数。
- 使用一种系统性地改变贝尔曼更新的框架,以模拟多种非理性类型,包括短视和双曲贴现。
- 在三个环境(随机MDP、网格世界和自动驾驶领域)中,使用对数损失作为指标评估奖励推理性能。
- 在不同人类模型下比较推理性能:噪声理性(Boltzmann)、短视规划者和正确指定的非理性模型。
- 通过测量策略与奖励之间的互信息,从理论上解释为何非理性行为能提升信息量。
- 通过参数和类型误设的消融研究,测试非理性行为建模的鲁棒性。
实验结果
研究问题
- RQ1将非理性人类建模为噪声理性行为,是否会导致比使用正确非理性模型更差的奖励推理性能?
- RQ2当非理性人类行为被正确建模时,是否能比完全理性的行为提供更多的奖励信息?
- RQ3在不同非理性类型下,行为与奖励之间的互信息如何变化?
- RQ4对非理性行为的近似模型在多大程度上仍优于标准的噪声理性假设?
- RQ5当机器人正确建模人类非理性行为时,奖励推理性能的上限是什么?
主要发现
- 即使人类行为完全理性,正确建模人类非理性行为,其奖励推理性能也优于将人类建模为理性行为。
- 短视行为可提高策略与奖励之间的互信息,使其在本质上比理性行为更具信息量,后者在某些情况下对奖励变化保持不变。
- 当机器人错误地假设人类行为为理性(例如Boltzmann理性)时,推理性能可能显著劣于使用先验信息,尤其是在存在系统性偏差时。
- 即使参数错误,只要建模了正确的非理性类型(例如短视),其推理性能仍优于假设为噪声理性。
- 在自动驾驶领域中,短视人类行为的奖励推理能力强于理性行为,证实了在高维、连续环境中的结果。
- 理论分析证明,某些非理性行为可使行为与奖励之间的互信息任意增加,从而使其信息量超过理性行为。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。