[Paper Review] Exploring the Efficacy of ChatGPT in Analyzing Student Teamwork Feedback with an Existing Taxonomy
The paper evaluates ChatGPT’s ability to label student teamwork comments according to an existing taxonomy and to self-assess its labeling accuracy, finding high alignment with human labels.
Teamwork is a critical component of many academic and professional settings. In those contexts, feedback between team members is an important element to facilitate successful and sustainable teamwork. However, in the classroom, as the number of teams and team members and frequency of evaluation increase, the volume of comments can become overwhelming for an instructor to read and track, making it difficult to identify patterns and areas for student improvement. To address this challenge, we explored the use of generative AI models, specifically ChatGPT, to analyze student comments in team based learning contexts. Our study aimed to evaluate ChatGPT's ability to accurately identify topics in student comments based on an existing framework consisting of positive and negative comments. Our results suggest that ChatGPT can achieve over 90\% accuracy in labeling student comments, providing a potentially valuable tool for analyzing feedback in team projects. This study contributes to the growing body of research on the use of AI models in educational contexts and highlights the potential of ChatGPT for facilitating analysis of student comments.
Motivation & Objective
- Motivate the use of teamwork feedback analysis in education and demonstrate scalability of feedback review in large classes.
- Evaluate ChatGPT’s ability to classify student comments into predefined topics from an existing taxonomy.
- Assess ChatGPT’s self-evaluated accuracy against human raters to gauge reliability in labeling.
- Explore practical implications and limitations for using generative AI in qualitative analysis of educational feedback.
Proposed method
- Use an archived, de-identified set of 200 student comments from an undergraduate course as test data.
- Apply zero-shot prompts to ChatGPT-3.5-turbo to identify topics from a provided taxonomy (positive and negative comments).
- Prompt ChatGPT to return topics per comment in a table format with original comment IDs and labeled topics.
- Assess labeling accuracy by having human researchers rate ChatGPT’s labels on a three-point scale (accurate, unclear, inaccurate).
- Prompt the model to perform an accuracy check by rating its own labeling accuracy on a 1–10 ordinal scale against human judgments.
Experimental results
Research questions
- RQ1RQ1: How well does the instruction-tuned GPT-3.5 model match human labels when classifying student feedback into taxonomy-based categories?
- RQ2RQ2: To what extent does the model’s self-rated labeling accuracy correspond to human evaluations on an ordinal scale?
Key findings
- ChatGPT achieved about 85% fully accurate labeling according to human raters for the 200 comments analyzed.
- The model produced 282 labels across comments, indicating some comments received multiple labels.
- The most common mislabel was Attended group meetings, reflecting occasional labeling errors and potential defaulting to the first option.
- Overall, the model demonstrated the ability to capture both positive and constructive feedback and to identify semantic meanings beyond exact wording.
- The study suggests that pre-trained generative models can qualitatively analyze open-ended feedback with high alignment to human labels without fine-tuning, though sentiment and label choice issues remain.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.