Skip to main content
QUICK REVIEW

[论文解读] Human-Robot Commensality: Bite Timing Prediction for Robot-Assisted Feeding in Groups

Jan Ondras, Abrar Anwar|arXiv (Cornell University)|Jul 7, 2022
Social Robot Interaction and HRI被引用 4
一句话总结

本文提出 SoNNET,一种多模态深度学习模型,通过分析人际社交线索,预测在群体聚餐场景中机器人辅助进餐的社交适宜进餐时机。基于一个包含30组三人聚餐的新型人际共餐数据集(HHCD),该模型利用音频、视频和注视信号推断最佳进餐时刻,在用户研究中显著提升了社交融合度与用户偏好,优于人工或固定间隔触发方式。

ABSTRACT

We develop data-driven models to predict when a robot should feed during social dining scenarios. Being able to eat independently with friends and family is considered one of the most memorable and important activities for people with mobility limitations. While existing robotic systems for feeding people with mobility limitations focus on solitary dining, commensality, the act of eating together, is often the practice of choice. Sharing meals with others introduces the problem of socially appropriate bite timing for a robot, i.e. the appropriate timing for the robot to feed without disrupting the social dynamics of a shared meal. Our key insight is that bite timing strategies that take into account the delicate balance of social cues can lead to seamless interactions during robot-assisted feeding in a social dining scenario. We approach this problem by collecting a Human-Human Commensality Dataset (HHCD) containing 30 groups of three people eating together. We use this dataset to analyze human-human commensality behaviors and develop bite timing prediction models in social dining scenarios. We also transfer these models to human-robot commensality scenarios. Our user studies show that prediction improves when our algorithm uses multimodal social signaling cues between diners to model bite timing. The HHCD dataset, videos of user studies, and code are available at https://emprise.cs.cornell.edu/hrcom/

研究动机与目标

  • 为解决机器人进餐系统在群体聚餐场景中缺乏社交意识的进餐时机问题。
  • 开发一种数据驱动模型,预测机器人在不破坏社交动态的前提下应何时为用户进餐。
  • 将从人际互动中学到的进餐时机策略成功迁移至人机共餐场景。
  • 在真实世界三人用户研究中评估模型性能与社交可接受性。

提出的方法

  • 收集了包含30组参与者共享聚餐的多视角RGBD视频与定向音频的人际共餐数据集(HHCD)。
  • 开发了SoNNET,一种社交小口进食网络,通过在6秒时间窗口内融合多模态信号(注视、言语、面部表情与手势),预测进餐意图。
  • 在人际互动上训练模型,以学习群体聚餐中的自然社交节奏与协调模式。
  • 将训练好的SoNNET模型迁移至人机进餐场景,在实体机器人上实现实时进餐时机决策。
  • 以口部张开检测触发作为基线,与人工和固定间隔触发在用户研究中进行性能对比。
  • 开展10组三人小组的用户研究,评估社交可接受性、对话流畅度与用户偏好。

实验结果

研究问题

  • RQ1机器人如何在共享聚餐过程中推断出社交适宜的进餐时机,而不干扰对话或群体动态?
  • RQ2在群体聚餐场景中,哪些多模态社交线索最能预测人类的进餐时机?
  • RQ3从人际互动中学到的进餐时机策略能否成功迁移至人机进餐场景?
  • RQ4与人工或固定间隔触发相比,基于模型的进餐时机在社交可接受性与用户体验方面表现如何?
  • RQ5用餐环境中的哪些因素(如机器人位置、语音提示、移动行为)最影响感知到的社交干扰?

主要发现

  • 在用户研究中,SoNNET模型显著优于人工和固定间隔触发方式,参与者认为其更自然且干扰更小。
  • 多模态社交线索(包括注视、言语停顿与面部表情)提升了进餐时机预测的准确性,并促进了更流畅的社交互动。
  • 70%的机器人用户认为在口部张开检测期间的语音提示具有干扰性,表明需要更安静或情境感知的提示机制。
  • 共同进餐者报告称,机器人的移动与语音打断会干扰对话流,尤其在时机不准或视线受阻时。
  • 尽管存在干扰,70%的用户与共同进餐者表示,当使用SoNNET时,机器人并未显著中断对话,许多人形容机器人是“一个酷酷的第四位共餐者”。
  • 数据集与代码已公开发布于 https://emprise.cs.cornell.edu/hrcom/,为未来社交意识机器人进餐研究提供支持。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。