[论文解读] Exploring the Efficacy of ChatGPT in Analyzing Student Teamwork Feedback with an Existing Taxonomy
这篇论文评估了 ChatGPT 将学生团队合作评论按照现有分类法进行标注的能力,并自我评估其标注准确性,发现与人工标签高度对齐。
Teamwork is a critical component of many academic and professional settings. In those contexts, feedback between team members is an important element to facilitate successful and sustainable teamwork. However, in the classroom, as the number of teams and team members and frequency of evaluation increase, the volume of comments can become overwhelming for an instructor to read and track, making it difficult to identify patterns and areas for student improvement. To address this challenge, we explored the use of generative AI models, specifically ChatGPT, to analyze student comments in team based learning contexts. Our study aimed to evaluate ChatGPT's ability to accurately identify topics in student comments based on an existing framework consisting of positive and negative comments. Our results suggest that ChatGPT can achieve over 90\% accuracy in labeling student comments, providing a potentially valuable tool for analyzing feedback in team projects. This study contributes to the growing body of research on the use of AI models in educational contexts and highlights the potential of ChatGPT for facilitating analysis of student comments.
研究动机与目标
- 在教育中激励对团队合作反馈进行分析的应用,并展示在大班级中的反馈评审的可扩展性。
- 评估 ChatGPT 将学生评论分类为来自现有分类法的预定义主题的能力。
- 评估 ChatGPT 的自评准确性相对于人工评估者,以衡量在标注中的可靠性。
- 探讨在教育反馈的定性分析中使用生成式AI的实际含义和局限性。
提出的方法
- 使用来自本科课程的经归档、去标识的200条学生评论作为测试数据。
- 对 ChatGPT-3.5-turbo 应用零-shot 提示,以从提供的分类法中识别主题(正面和负面评论)。
- 提示 ChatGPT 以表格格式返回每条评论的主题,包含原始评论ID和标注主题。
- 通过让人工研究者用三分制评估 ChatGPT 的标签(准确、含糊、不准确)来评估标注准确性。
- 要求模型执行一个准确性检查,按1–10的序数尺度对比人工判断对其自身标注准确性的评分。
实验结果
研究问题
- RQ1RQ1:当将学生反馈分类为基于分类法的类别时,指令微调的 GPT-3.5 模型与人类标签的匹配程度如何?
- RQ2RQ2:模型自评的标注准确性在多大程度上与人类在序数尺度上的评估相吻合?
主要发现
- 根据人工评审,对分析的200条评论,ChatGPT 实现约85%的完全准确标注。
- 模型在评论中生成了282个标签,表明某些评论获得了多个标签。
- 最常见的错误标注是“出席小组会议”,反映偶发的标注错误和潜在的默认使用第一个选项。
- 总体上,该模型展现了捕捉积极和建设性反馈的能力,并能识别超出原文措辞的语义含义。
- 研究表明,在不进行微调的情况下,预训练的生成模型可以对开放式反馈进行定性分析,与人工标签高度对齐,尽管情感和标签选择问题仍然存在。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。