[论文解读] Prediction and Localization of Student Engagement in the Wild
本文引入了一个新的‘真实场景’数据集,包含来自78名受试者的195段视频,总计16.5小时,标注了四个层次的参与度,以支持学生参与度预测的弱监督学习。该研究提出了一种深度多实例学习框架,利用面部、视线和身体动作线索定位具有参与度和无参与度的视频片段,为改进MOOC视频设计提供了洞见。
In this paper, we introduce a new dataset for student engagement detection and localization. Digital revolution has transformed the traditional teaching procedure and a result analysis of the student engagement in an e-learning environment would facilitate effective task accomplishment and learning. Well known social cues of engagement/disengagement can be inferred from facial expressions, body movements and gaze pattern. In this paper, student's response to various stimuli videos are recorded and important cues are extracted to estimate variations in engagement level. In this paper, we study the association of a subject's behavioral cues with his/her engagement level, as annotated by labelers. We then localize engaging/non-engaging parts in the stimuli videos using a deep multiple instance learning based framework, which can give useful insight into designing Massive Open Online Courses (MOOCs) video material. Recognizing the lack of any publicly available dataset in the domain of user engagement, a new `in the wild' dataset is created to study the subject engagement problem. The dataset contains 195 videos captured from 78 subjects which is about 16.5 hours of recording. We present detailed baseline results using different classifiers ranging from traditional machine learning to deep learning based approaches. The subject independent analysis is performed so that it can be generalized to new users. The problem of engagement prediction is modeled as a weakly supervised learning problem. The dataset is manually annotated by different labelers for four levels of engagement independently and the correlation studies between annotated and predicted labels of videos by different classifiers is reported. This dataset creation is an effort to facilitate research in various e-learning environments such as intelligent tutoring systems, MOOCs, and others.
研究动机与目标
- 为解决真实世界在线学习环境中学生参与度缺乏公开可用数据集的问题。
- 将学生参与度预测建模为使用视频级别标注的弱监督学习问题。
- 利用行为线索定位教育视频中具有参与度和无参与度的片段。
- 为在用户无关设置下评估参与度预测模型提供基准。
- 支持自适应在线学习系统(如MOOC和智能辅导系统)的开发。
提出的方法
- 在自然主义的在线学习环境中收集78名受试者的195段视频,记录面部表情、视线模式和身体动作。
- 通过多名标注者对每段视频进行四档参与度标注,以确保可靠性。
- 提取多模态行为线索(面部、视线和运动特征)作为参与度建模的输入。
- 应用深度多实例学习(MIL)框架,定位具有参与度和无参与度的视频片段。
- 在该数据集上训练并评估多种分类器——涵盖传统机器学习到深度学习方法。
- 进行用户无关的评估,以确保模型对未见用户的泛化能力。
实验结果
研究问题
- RQ1在真实世界在线学习视频中,多模态行为线索(面部、视线、身体)与人工标注的参与度水平之间有何相关性?
- RQ2仅使用视频级别标签,弱监督学习方法能否有效定位教育视频中的具有参与度和无参与度的片段?
- RQ3不同分类器(传统机器学习与深度学习)在该新数据集上预测参与度时的性能表现如何?
- RQ4多位标注者之间的参与度标注一致性如何?这种一致性对模型训练有何影响?
- RQ5所提出的框架在用户无关评估设置下是否能泛化到新用户?
主要发现
- 所提出的深度多实例学习框架在教育视频中成功定位具有参与度和无参与度的片段,其准确率优于基线模型。
- 数据集在不同标注者之间表现出中等到较高的组内一致性,验证了参与度标注的可靠性。
- 用户无关评估表明,该模型对新用户具有良好的泛化能力,这是实际部署的关键要求。
- 传统机器学习模型表现具有竞争力,但深度学习方法在定位准确率方面表现更优。
- 引入视线和身体动作特征显著提升了参与度预测性能,相较于仅使用面部特征有明显改进。
- 该数据集为未来参与度检测研究提供了稳健的基准,尤其适用于MOOC和智能辅导系统。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。