[论文解读] Mining MOOC Clickstreams: On the Relationship Between Learner Video-Watching Behavior and Performance.
本文提出了两种点击流框架——基于事件的和基于位置的——以建模MOOC学习者观看视频的行为,并预测其在测验题上的表现。研究发现,如反思和复习等重复性行为模式与首次作答正确(CFA)显著相关,且基于位置的模型在CFA预测准确率和F1分数上均有提升,尤其在数据稀缺的场景(如短课程或早期检测)中更具实用性。
We study student behavior and performance in two Massive Open Online Courses (MOOCs). In doing so, we present two frameworks by which video-watching clickstreams can be represented: one based on the sequence of events created, and another on the sequence of positions visited. With the event-based framework, we extract recurring subsequences of student behavior, which contain fundamental characteris- tics such as reflecting (i.e., repeatedly playing and pausing) and revising (i.e., plays and skip backs). We find that some of these behaviors are significantly associated with whether a user will be Correct on First Attempt (CFA) or not in answering quiz questions. With the position-based framework, we then devise models for performance. In evaluating these through CFA prediction, we find that three of them can substantially improve prediction quality in terms of accuracy and F1, which underlines the ability to relate behavior to performance. Since our prediction considers videos individually, these benefits also suggest that our models are useful in situations where there is limited training data, e.g., for early detection or in short courses.
研究动机与目标
- 理解MOOC中学习者观看视频的行为与其在测验题上表现之间的关系。
- 解决在训练数据有限的情况下(特别是在短课程或早期阶段课程中)预测学习者表现的挑战。
- 开发能够捕捉点击流数据中有意义行为模式的行为表征框架,用于表现建模。
- 评估从点击流中提取的行为特征是否能够增强对首次作答正确(CFA)结果的预测能力。
提出的方法
- 基于事件的框架将点击流数据建模为一系列操作(如播放、暂停、快进)的序列,从中提取代表反思和复习等行为的重复子序列。
- 基于位置的框架将视频观看行为表示为播放位置的序列,从而能够建模学习者进度与重复观看模式。
- 利用序列挖掘技术识别行为子序列,以检测如重复播放和倒退导航等模式。
- 基于位置的表示构建表现预测模型,整合行为特征以预测CFA结果。
- 采用标准指标(准确率与F1)在CFA预测任务上评估模型性能,以衡量其相对于基线方法的改进程度。
- 该方法设计为在训练数据有限时依然有效,支持早期检测,并适用于短课程场景。
实验结果
研究问题
- RQ1哪些重复出现的视频观看行为(如反思或复习)与首次作答正确(CFA)表现显著相关?
- RQ2如何有效表示点击流数据,以捕捉MOOC视频消费中学习者有意义的行为?
- RQ3与基线模型相比,基于行为的模型在多大程度上能提升CFA结果的预测能力?
- RQ4在低数据场景(如早期检测或短课程)中,所提出的模型是否能保持高性能?
主要发现
- 如反复播放并暂停(反思)和播放后倒退(复习)等重复性行为模式,与学习者首次作答是否正确显著相关。
- 基于位置的性能预测模型中有三种在准确率和F1分数上均较基线模型实现显著提升。
- 预测质量的提升表明,视频观看行为中包含对表现估计具有意义的信号。
- 模型在低数据设置下的有效性表明其在早期表现检测以及短课程或新兴MOOC中的应用潜力巨大。
- 基于事件的框架成功识别出与学习成果相关的根本性行为特征。
- 基于位置的建模方法实现了精确到单个视频级别的表现预测,提升了可解释性与实际应用价值。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。