[论文解读] Learned human-agent decision-making, communication and joint action in a virtual reality environment
本研究引入了一种虚拟现实觅食环境,以探究人类与智能体的协同行动,其中人类飞行员与AI副驾驶通过预测学习和基于策略的通信实现共同适应。副驾驶通过上下文Bandit策略学习发出音频提示,随时间推移提升共享奖励;结果表明,夜间觅食能力得到增强,且出现协同行为,尽管初期因提示时机和学习延迟导致出现错误。
Humans make decisions and act alongside other humans to pursue both short-term and long-term goals. As a result of ongoing progress in areas such as computing science and automation, humans now also interact with non-human agents of varying complexity as part of their day-to-day activities; substantial work is being done to integrate increasingly intelligent machine agents into human work and play. With increases in the cognitive, sensory, and motor capacity of these agents, intelligent machinery for human assistance can now reasonably be considered to engage in joint action with humans---i.e., two or more agents adapting their behaviour and their understanding of each other so as to progress in shared objectives or goals. The mechanisms, conditions, and opportunities for skillful joint action in human-machine partnerships is of great interest to multiple communities. Despite this, human-machine joint action is as yet under-explored, especially in cases where a human and an intelligent machine interact in a persistent way during the course of real-time, daily-life experience. In this work, we contribute a virtual reality environment wherein a human and an agent can adapt their predictions, their actions, and their communication so as to pursue a simple foraging task. In a case study with a single participant, we provide an example of human-agent coordination and decision-making involving prediction learning on the part of the human and the machine agent, and control learning on the part of the machine agent wherein audio communication signals are used to cue its human partner in service of acquiring shared reward. These comparisons suggest the utility of studying human-machine coordination in a virtual reality environment, and identify further research that will expand our understanding of persistent human-machine joint action.
研究动机与目标
- 在自然、沉浸式的VR环境中,研究持久的、实时的人机协同行动。
- 研究人类与学习型智能体如何通过预测学习和通信实现共同适应。
- 评估不同AI副驾驶架构(Pavlovian提示 vs. 上下文Bandit策略学习)对人类决策与协调的影响。
- 探索在持续、动态交互中人类-智能体伙伴关系的涌现式通信与技能迁移。
提出的方法
- 设计了一项虚拟现实觅食任务,其中六种水果在昼夜光照之间循环变化,要求人类与智能体实时协作。
- 人类飞行员通过手柄控制器进行果实采摘,而AI副驾驶则基于学习到的价值预测来判断果实成熟度并发出音频提示。
- 副驾驶使用上下文Bandit算法,更新其用于提示的随机策略,强化能带来更高共享奖励的行为。
- 通过累计得分、提示响应时间以及预测值变化($\Delta V(h,s)$)追踪人类飞行员行为,反映学习与教学信号。
- 测试了三种实验条件:无副驾驶(NoCP)、Pavlovian提示(Pav)和策略学习副驾驶(Bandit),后者支持自适应提示。
- 副驾驶的策略基于奖励反馈进行更新,信用分配依赖于飞行员的行为与时机,从而模拟实时协同行动。
实验结果
研究问题
- RQ1AI副驾驶的存在如何影响人类在持久、实时VR环境中的觅食行为?
- RQ2基于策略学习的副驾驶能否通过自适应、情境敏感的通信提升共享奖励?
- RQ3时机与信用分配在有效的人机协同行动与通信中扮演何种角色?
- RQ4不同副驾驶架构(Pavlovian vs. 上下文Bandit)如何影响人类学习与协调?
- RQ5人类与智能体的预测与控制策略在共享任务中通过相互学习能实现多大程度的共同适应?
主要发现
- 基于策略学习的副驾驶(Bandit)在夜间觅食活动显著多于无副驾驶条件,表明其对低能见度条件的适应能力更强。
- 尽管觅食活动增加,但飞行员在两种副驾驶条件下均出现更多错误,尤其是在夜间和不熟悉果实位置时,表明对提示解释存在学习延迟。
- 在排除负分事件后的总分显示,与Bandit副驾驶合作可在双方学会最优协调后实现更高的有效性能。
- 通过预测值变化($\Delta V(h,s)$)衡量的教学互动表明,飞行员随时间适应了副驾驶的提示,表明存在相互学习。
- Bandit条件的提示策略比Pavlovian方法更有效且更渐进,支持自适应、情境敏感通信的必要性。
- 时间延迟与信用分配至关重要:提示错位会导致错误,凸显了在人机协同行动中精确时间对齐的重要性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。