Skip to main content
QUICK REVIEW

[論文レビュー] Prediction and Localization of Student Engagement in the Wild

Aamir Mustafa, Amanjot Kaur|arXiv (Cornell University)|Apr 3, 2018
Online Learning and Analytics被引用数 18
ひとこと要約

本論文は、78名の被験者から収集した195本の動画(合計16.5時間)からなる、実世界のeラーニング環境を想定した新しい「in the wild」データセットを紹介している。このデータセットは、4段階の関与度をラベル付けしており、学生の関与度予測の弱教師付き学習を可能にする。本研究では、顔貌、視線、身体の動きの特徴を用いて、関与度の高いおよび低い動画セグメントを局所化するための深層多重インスタンス学習フレームワークを提案しており、MOOCの動画設計の改善に役立つ知見を提供する。

ABSTRACT

In this paper, we introduce a new dataset for student engagement detection and localization. Digital revolution has transformed the traditional teaching procedure and a result analysis of the student engagement in an e-learning environment would facilitate effective task accomplishment and learning. Well known social cues of engagement/disengagement can be inferred from facial expressions, body movements and gaze pattern. In this paper, student's response to various stimuli videos are recorded and important cues are extracted to estimate variations in engagement level. In this paper, we study the association of a subject's behavioral cues with his/her engagement level, as annotated by labelers. We then localize engaging/non-engaging parts in the stimuli videos using a deep multiple instance learning based framework, which can give useful insight into designing Massive Open Online Courses (MOOCs) video material. Recognizing the lack of any publicly available dataset in the domain of user engagement, a new `in the wild' dataset is created to study the subject engagement problem. The dataset contains 195 videos captured from 78 subjects which is about 16.5 hours of recording. We present detailed baseline results using different classifiers ranging from traditional machine learning to deep learning based approaches. The subject independent analysis is performed so that it can be generalized to new users. The problem of engagement prediction is modeled as a weakly supervised learning problem. The dataset is manually annotated by different labelers for four levels of engagement independently and the correlation studies between annotated and predicted labels of videos by different classifiers is reported. This dataset creation is an effort to facilitate research in various e-learning environments such as intelligent tutoring systems, MOOCs, and others.

研究の動機と目的

  • 実世界のeラーニング環境における学生の関与度を測定するための公開データセットが不足しているという問題に取り組む。
  • 動画レベルのラベルのみを用いて、弱教師付き学習の枠組みで学生の関与度予測をモデル化する。
  • 行動的特徴(顔貌、視線、動き)を用いて、教育動画内の関与度の高いおよび低いセグメントを局所化する。
  • 被験者独立の設定における関与度予測モデルの評価のためのベンチマークを提供する。
  • MOOCやインテリジェントチューティングシステムを含む、適応型eラーニングシステムの開発を支援する。

提案手法

  • 自然なeラーニング環境で、78名の被験者から195本の動画を収集し、顔貌の変化、視線のパターン、身体の動きを記録する。
  • 複数のラベルラーによるラベル付けを用いて、各動画に4段階の関与度を付与し、信頼性を確保する。
  • 関与度モデリングの入力として、マルチモーダルな行動的特徴(顔貌、視線、動きの特徴)を抽出する。
  • 関与度の高いおよび低い動画セグメントを局所化するために、深層多重インスタンス学習(MIL)フレームワークを適用する。
  • 本データセット上で、従来の機械学習から深層学習まで多様な分類器を訓練および評価する。
  • 一般化性能を検証するため、被験者独立の評価を実施する。

実験結果

リサーチクエスチョン

  • RQ1実世界のeラーニング動画において、顔貌、視線、身体の動きといったマルチモーダルな行動的特徴は、人間ラベルラーが付与した関与度レベルとどの程度相関しているか?
  • RQ2動画レベルのラベルのみを用いて、弱教師付き学習アプローチが教育動画内の関与度の高いおよび低いセグメントを効果的に局所化できるか?
  • RQ3本データセット上での、従来の機械学習と深層学習の分類器の性能は、関与度予測においてどのように異なるか?
  • RQ4複数のラベルラーによる関与度ラベル付けの間にはどの程度の合意が得られており、それがモデル学習にどのように影響するか?
  • RQ5提案されたフレームワークは、被験者独立の評価設定において、新しいユーザーに対しても一般化可能か?

主な発見

  • 提案された深層多重インスタンス学習フレームワークは、ベースラインモデルよりも高い精度で教育動画内の関与度の高いおよび低いセグメントを効果的に局所化した。
  • ラベルラー間の関与度ラベル付けには中程度から高い一貫性が見られ、関与度ラベル付けの信頼性が裏付けられた。
  • 被験者独立の評価から、モデルが新しいユーザーに対しても良好に一般化することが示された。これは、実世界への導入に不可欠な要件である。
  • 従来の機械学習モデルは競争力のある性能を示したが、深層学習アプローチはより優れた局所化精度を達成した。
  • 視線と身体の動きの特徴を組み込むことで、顔貌特徴のみを用いた場合に比べ、関与度予測の性能が顕著に向上した。
  • 本データセットは、将来的な関与度検出研究、特にMOOCやインテリジェントチューティングシステム分野における強固なベンチマークを提供する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。