Skip to main content
QUICK REVIEW

[论文解读] When Humans Aren't Optimal: Robots that Collaborate with Risk-Aware Humans

Minae Kwon, Erdem Bıyık|arXiv (Cornell University)|Jan 13, 2020
Decision-Making and Behavioral Economics参考文献 46被引用 4
一句话总结

本文提出了一种基于累积前景理论的风险感知人类模型,以在不确定性条件下提升人机协作性能。通过用风险敏感的框架替代传统的噪声理性模型,机器人能够更好地预测在高风险场景下人类的次优行为,从而在自动驾驶和叠杯任务中实现更安全、更高效的交互。

ABSTRACT

In order to collaborate safely and efficiently, robots need to anticipate how their human partners will behave. Some of today's robots model humans as if they were also robots, and assume users are always optimal. Other robots account for human limitations, and relax this assumption so that the human is noisily rational. Both of these models make sense when the human receives deterministic rewards: i.e., gaining either $100 or $130 with certainty. But in real world scenarios, rewards are rarely deterministic. Instead, we must make choices subject to risk and uncertainty--and in these settings, humans exhibit a cognitive bias towards suboptimal behavior. For example, when deciding between gaining $100 with certainty or $130 only 80% of the time, people tend to make the risk-averse choice--even though it leads to a lower expected gain! In this paper, we adopt a well-known Risk-Aware human model from behavioral economics called Cumulative Prospect Theory and enable robots to leverage this model during human-robot interaction (HRI). In our user studies, we offer supporting evidence that the Risk-Aware model more accurately predicts suboptimal human behavior. We find that this increased modeling accuracy results in safer and more efficient human-robot collaboration. Overall, we extend existing rational human models so that collaborative robots can anticipate and plan around suboptimal human behavior during HRI.

研究动机与目标

  • 为解决现有机器人假设人类完全理性或存在噪声理性的局限性,此类假设在真实世界的风险与不确定性情境下失效。
  • 将人类决策建模为风险感知而非纯粹理性,以反映在不确定性下存在的认知偏差,如风险规避或风险寻求。
  • 通过将行为经济学原理整合到人机交互(HRI)模型中,改进机器人规划与协作能力。
  • 通过实证验证,风险感知模型在预测人类行为及提升协作结果方面优于噪声理性模型。

提出的方法

  • 采用累积前景理论(CPT)作为核心人类决策模型,以捕捉概率与结果的非线性加权。
  • 构建一种心智理论(ToM)框架,使机器人能够基于对奖励与概率的感知,建模人类的风险敏感偏好。
  • 将CPT整合到概率推理系统中,以在风险情境下预测人类行为,取代标准的期望效用最大化,转而使用前景价值函数。
  • 使用相同的人类示范数据,训练并对比两种机器人模型:噪声理性基线模型与风险感知模型。
  • 利用预测的人类行为指导机器人规划,减少干扰,提升协作场景下的任务效率。
  • 在模拟自动驾驶和真实的叠杯任务中开展用户研究,以评估模型的预测准确性和协作质量。

实验结果

研究问题

  • RQ1基于累积前景理论的风险感知人类模型是否能比噪声理性模型更准确地预测不确定、高风险情境下的人类行为?
  • RQ2在哪些类型的决策情境中,风险感知建模对提升机器人预测与协作能力最为关键?
  • RQ3与传统理性模型相比,使用风险感知模型是否能带来更安全、更高效的协作?
  • RQ4人类用户如何感知并回应能够预判其风险规避或风险寻求倾向的机器人?

主要发现

  • 基于累积前景理论的风险感知模型在自动驾驶模拟和真实叠杯任务中,显著优于噪声理性基线模型,对人类行为的预测表现更优。
  • 在叠杯任务中,参与者与风险感知机器人协作时完成任务更快,且报告的干扰更少,相比噪声理性机器人。
  • 参与者主观上更偏好风险感知机器人,认为其更高效且更符合自身意图,虽偏好程度边际显著但为正(p < .07)。
  • 用户报告称,噪声理性机器人常因试图抓取用户正在伸手的同一杯子而造成干扰,从而增加任务耗时。
  • 风险感知机器人通过预判人类会避免高风险行为(如抓取不稳定的杯子),有效减少了干扰。
  • 在具有更长规划时域的网格世界实验中,风险感知模型在预测人类行为序列方面保持了比噪声理性模型更高的准确性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。