Skip to main content
QUICK REVIEW

[论文解读] Engagement Detection in Meetings

Maria Frank, Ghassem Tofighi|arXiv (Cornell University)|Aug 31, 2016
Personal Information Management and User Behavior参考文献 9被引用 9
一句话总结

本文提出了一种多模态框架,利用2D/3D成像、音频和运动数据实现实时群体会议参与度检测,将参与者分类为六种参与状态。该系统在分类‘行动意图’状态时达到83.36%的准确率,支持动态反馈以提升会议效率和参与者意识。

ABSTRACT

Group meetings are frequent business events aimed to develop and conduct project work, such as Big Room design and construction project meetings. To be effective in these meetings, participants need to have an engaged mental state. The mental state of participants however, is hidden from other participants, and thereby difficult to evaluate. Mental state is understood as an inner process of thinking and feeling, that is formed of a conglomerate of mental representations and propositional attitudes. There is a need to create transparency of these hidden states to understand, evaluate and influence them. Facilitators need to evaluate the meeting situation and adjust for higher engagement and productivity. This paper presents a framework that defines a spectrum of engagement states and an array of classifiers aimed to detect the engagement state of participants in real time. The Engagement Framework integrates multi-modal information from 2D and 3D imaging and sound. Engagement is detected and evaluated at participants and aggregated at group level. We use empirical data collected at the lab of Konica Minolta, Inc. to test initial applications of this framework. The paper presents examples of the tested engagement classifiers, which are based on research in psychology, communication, and human computer interaction. Their accuracy is illustrated in dyadic interaction for engagement detection. In closing we discuss the potential extension to complex group collaboration settings and future feedback implementations.

研究动机与目标

  • 开发一种实时框架,用于检测和分类群体会议中的参与状态,以提升协作与生产力。
  • 揭示隐藏的心理状态(尤其是参与度和行动意图)的透明度,以支持主持人调整会议动态。
  • 通过整合多模态输入(姿势、面部表情、语音、运动)扩展eRing原型,超越仅依赖身体运动的限制。
  • 在个体和群体层面实现反馈机制,以重新吸引注意力不集中的参与者或调整会议焦点。
  • 为复杂协作环境(如大空间设计会议)中的可扩展、基于传感器的参与度监测奠定基础。

提出的方法

  • 该框架基于行为和感知线索,将参与度分类为六种状态:倾听、观察、发言、行动、行动意图和脱离参与。
  • 实时处理来自2D/3D摄像机、音频和运动传感器的多模态数据流,提取手部速度、姿势和面部表情等特征。
  • 使用支持向量机(SVM)进行参与度状态分类,单帧处理时间低于10ms,支持每秒30帧的实时运行。
  • 系统使用动态参与度表来追踪参与者与房间内物体之间的潜在参与水平,支持基于意图的交互。
  • 分类器采用模块化设计且可扩展,未来可集成语音和生物特征数据等额外模态。
  • 初始实现聚焦于两人互动,并利用柯尼卡美能达实验室的实证数据训练和验证分类器。

实验结果

研究问题

  • RQ1如何利用多模态传感器数据有效检测和分类群体会议中的参与度?
  • RQ2在协作环境中,哪些行为和感知指标最可靠地反映参与度和行动意图?
  • RQ3通过反馈机制,实时参与度分类能否改善会议动态和参与者生产力?
  • RQ4所提出的多模态分类系统在实时检测参与度状态方面的准确性和效率如何?
  • RQ5在复杂环境(如大空间项目会议)中,自动化参与度反馈对会议主持和团队协作有何影响?

主要发现

  • 参与度分类器在检测两人互动中的‘行动意图’状态时,分类准确率达到83.36%。
  • 单帧处理时间低于10ms,证实了在每秒30帧的实时会议中部署的可行性。
  • 系统成功检测并聚合了个体和群体层面的参与度状态,支持动态反馈。
  • 当超过40%的参与者出现脱离参与状态时,可触发反馈机制,如公开提示(“团队已脱离参与”)或eRing的视觉信号。
  • 模块化设计支持未来扩展,可集成语音、面部表情和生物特征数据等额外模态以提升准确性。
  • 该框架支持操作意图反馈,使智能房间中的响应式物体能够基于方向性姿势和参与度水平,识别出哪位用户意图与哪个物体交互。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。