Skip to main content
QUICK REVIEW

[Paper Review] An Identification of Learners' Confusion through Language and Discourse Analysis

Thushari Atapattu, Katrina Falkner|arXiv (Cornell University)|Mar 8, 2019
Intelligent Tutoring Systems and Adaptive Learning37 references4 citations
TL;DR

This paper proposes a novel, language- and discourse-based approach to identify learner confusion in MOOCs, leveraging a custom linguistic feature set to classify confusion without relying on community metrics or physiological sensors. The method achieves strong cross-domain generalization, outperforming prior models by capturing individual-level affective states through textual analysis alone.

ABSTRACT

The substantial growth of online learning, in particular, Massively Open Online Courses (MOOCs), supports research into the development of better models for effective learning. Learner 'confusion' is among one of the identified aspects which impacts the overall learning process, and ultimately, course attrition. Confusion for a learner is an individual state of bewilderment and uncertainty of how to move forward. The majority of recent works neglect the 'individual' factor and measure the influence of community-related aspects (e.g. votes, views) for confusion classification. While this is a useful measure, as the popularity of one's post can indicate that many other students have similar confusion regarding course topics, these models neglect the personalised context, such as individual's affect or emotions. Certain physiological aspects (e.g. facial expressions, heart rate) have been utilised to classify confusion in small to medium classrooms. However, these techniques are challenging to adopt to MOOCs. To bridge this gap, we propose an approach solely based on language and discourse aspects of learners, which outperforms the previous models. We contribute through the development of a novel linguistic feature set that is predictive for confusion classification. We train the confusion classifier using one domain, successfully applying it across other domains.

Motivation & Objective

  • To address the gap in existing confusion detection models that overlook individual learner affect and emotions in MOOCs.
  • To develop a scalable, language-only approach to detect confusion that does not depend on community-level signals like votes or views.
  • To create a novel set of linguistic and discourse features predictive of confusion in learner-generated text.
  • To enable effective confusion classification across different course domains using a single-train, multi-domain generalization strategy.

Proposed method

  • The authors design a custom linguistic feature set capturing syntactic complexity, lexical diversity, hedging, and discourse markers from learner posts.
  • They extract textual features from learner forum posts in MOOCs, focusing on expressions of uncertainty, hesitation, and disfluency.
  • A supervised machine learning classifier is trained on one course domain using the proposed linguistic features to predict confusion states.
  • The model is evaluated for zero-shot cross-domain generalization, applying a model trained on one course to predict confusion in unrelated courses.
  • The approach avoids reliance on physiological sensors or community engagement metrics, focusing solely on linguistic cues.
  • The framework is validated across multiple MOOCs, demonstrating robustness and transferability of the linguistic features.

Experimental results

Research questions

  • RQ1Can linguistic and discourse features alone effectively identify learner confusion in MOOCs without community or physiological data?
  • RQ2How well does a confusion classifier trained on one course domain generalize to other, unrelated course domains?
  • RQ3What specific linguistic features are most predictive of learner confusion in open online learning environments?
  • RQ4To what extent does the proposed model outperform existing confusion detection methods that rely on community signals or sensor data?

Key findings

  • The proposed linguistic feature set significantly improves confusion detection performance compared to baseline models that rely on community metrics.
  • The model achieves strong cross-domain generalization, successfully applying a single-train model to predict confusion in diverse course topics.
  • Learner posts exhibiting hesitation, lexical simplification, and discourse uncertainty are strong indicators of confusion.
  • The approach outperforms existing methods in both domain-specific and zero-shot cross-domain settings, demonstrating the value of individualized textual analysis.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.