Skip to main content
QUICK REVIEW

[论文解读] Face valuing: Training user interfaces with facial expressions and reinforcement learning

Vivek Veeriah, Patrick M. Pilarski|arXiv (Cornell University)|Jun 9, 2016
Emotion and Mood Recognition参考文献 19被引用 21
一句话总结

本文提出了一种名为'面部估值'(face valuing)的强化学习方法,使智能体能够从面部表情推断用户偏好,从而减少对显式反馈的依赖。通过学习将面部特征映射到预期奖励的价值函数,智能体在抓握选择任务中能更快适应用户偏好,显著减少所需的人工修正——展示了在人机协作中一种可扩展、低认知负荷的方法。

ABSTRACT

An important application of interactive machine learning is extending or amplifying the cognitive and physical capabilities of a human. To accomplish this, machines need to learn about their human users' intentions and adapt to their preferences. In most current research, a user has conveyed preferences to a machine using explicit corrective or instructive feedback; explicit feedback imposes a cognitive load on the user and is expensive in terms of human effort. The primary objective of the current work is to demonstrate that a learning agent can reduce the amount of explicit feedback required for adapting to the user's preferences pertaining to a task by learning to perceive a value of its behavior from the human user, particularly from the user's facial expressions---we call this face valuing. We empirically evaluate face valuing on a grip selection task. Our preliminary results suggest that an agent can quickly adapt to a user's changing preferences with minimal explicit feedback by learning a value function that maps facial features extracted from a camera image to expected future reward. We believe that an agent learning to perceive a value from the body language of its human user is complementary to existing interactive machine learning approaches and will help in creating successful human-machine interactive applications.

研究动机与目标

  • 减少人机交互中显式反馈的认知与操作负担。
  • 使智能体能够通过面部表情等隐式信号学习用户偏好。
  • 开发一种可扩展、实时适应用户偏好变化而无需重新训练的方法。
  • 探索将面部表情作为价值信号而非控制输入的潜力。
  • 通过一种低努力的反馈机制,补充现有交互式机器学习方法。

提出的方法

  • 智能体使用时序差分学习训练一个价值函数,将从摄像头图像中提取的面部特征映射到预期未来奖励。
  • 面部特征实时处理以推断用户满意度,该结果作为策略学习的奖励信号。
  • 系统无需显式奖励通道,完全依赖隐式面部反馈运行。
  • 在模拟抓握选择任务中,一个目标导向的智能体基于面部表情学习选择最优抓握方式。
  • 智能体学会在获得肯定表情后再执行动作,从而最小化错误操作。
  • 该方法与任务无关,可泛化至其他形式的身体语言。

实验结果

研究问题

  • RQ1面部表情能否作为强化学习中用户满意度的可靠代理信号?
  • RQ2面部估值在多大程度上可减少交互式学习中对显式人工反馈的需求?
  • RQ3仅使用面部反馈时,智能体多快能适应用户偏好的变化?
  • RQ4基于面部特征学习的价值函数能否优于传统基于反馈的方法?
  • RQ5在具有新物体的动态、开放式环境中,该系统表现如何?

主要发现

  • 与仅依赖显式反馈的传统智能体相比,面部估值智能体在适应用户偏好方面显著更快。
  • 在无限物体场景中,面部估值智能体所需用户修正明显更少——在困难条件下反馈减少50-70%。
  • 智能体学会在获得肯定面部表情后再执行动作,从而减少错误并提高效率。
  • 实证结果表明,面部估值智能体能更一致地完成任务,且用户干预更少。
  • 即使用户偏好随时间变化,系统也能在不重新训练的情况下实现稳健适应。
  • 该方法在模拟环境中表现有效,预计在真实世界机器人应用中具有更好的可扩展性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。