Skip to main content
QUICK REVIEW

[论文解读] Is ChatGPT a Good Teacher Coach? Measuring Zero-Shot Performance For Scoring and Providing Actionable Insights on Classroom Instruction

Rose E. Wang, Dorottya Demszky|arXiv (Cornell University)|Jun 5, 2023
Online Learning and Analytics被引用 4
一句话总结

本文评估了ChatGPT在零样本提示下作为自动化教师教练的潜力,通过三项任务检验其表现:使用观察量规(CLASS和MQI)对课堂实录进行评分,识别教学亮点与改进机会,以及生成可操作的反馈以促进学生推理。尽管其回应听起来相关,但ChatGPT频繁建议教师已实施的行动(82%的建议如此),且其反馈缺乏洞察力与新颖性,表明在生成真实、原创且具有教学意义的教练反馈方面仍有显著改进空间。

ABSTRACT

Coaching, which involves classroom observation and expert feedback, is a widespread and fundamental part of teacher training. However, the majority of teachers do not have access to consistent, high quality coaching due to limited resources and access to expertise. We explore whether generative AI could become a cost-effective complement to expert feedback by serving as an automated teacher coach. In doing so, we propose three teacher coaching tasks for generative AI: (A) scoring transcript segments based on classroom observation instruments, (B) identifying highlights and missed opportunities for good instructional strategies, and (C) providing actionable suggestions for eliciting more student reasoning. We recruit expert math teachers to evaluate the zero-shot performance of ChatGPT on each of these tasks for elementary math classroom transcripts. Our results reveal that ChatGPT generates responses that are relevant to improving instruction, but they are often not novel or insightful. For example, 82% of the model's suggestions point to places in the transcript where the teacher is already implementing that suggestion. Our work highlights the challenges of producing insightful, novel and truthful feedback for teachers while paving the way for future research to address these obstacles and improve the capacity of generative AI to coach teachers.

研究动机与目标

  • 探究类似ChatGPT的生成式AI是否可作为专家教师教练的可扩展、低成本替代方案。
  • 评估ChatGPT在三项核心教师教练任务上的零样本表现:评分、识别教学优势与不足,以及生成可操作反馈。
  • 与专家人类评分对比,评估AI生成反馈的相关性、忠实度、洞察力与新颖性。
  • 识别当前大语言模型在真实课堂情境中提供高质量、具有教学意义教练反馈方面的局限性。

提出的方法

  • 使用标注过的小学数学课堂实录数据集(NCTE数据集),并以CLASS和MQI观察工具进行标注。
  • 采用零样本提示方法,评估ChatGPT在三项任务上的表现:(A) 使用量规条目对实录片段进行评分,(B) 识别亮点与改进机会,(C) 生成促进学生推理的建议。
  • 收集专家数学教师对模型输出质量的人工评估,评估标准包括相关性、忠实度、洞察力与可操作性。
  • 使用加权Cohen’s kappa计算标注者间一致性,相关性为0.23,忠实度为0.36,洞察力为0.37。
  • 将模型生成的评分与人工标注的量规评分进行比较,评估相关性,发现一致性较低。
  • 通过检查建议是否提及实录中已存在的行为,分析模型输出的冗余性。
Figure 1: Setup for the automated feedback task. Our work proposes three teacher coaching tasks. Task A is to score a transcript segment for items derived from classroom observation instruments; for instance, CLPC, CLBM, and CLINSTD are CLASS observation items, and EXPL, REMED, LANGIMP, SMQR are MQI
Figure 1: Setup for the automated feedback task. Our work proposes three teacher coaching tasks. Task A is to score a transcript segment for items derived from classroom observation instruments; for instance, CLPC, CLBM, and CLINSTD are CLASS observation items, and EXPL, REMED, LANGIMP, SMQR are MQI

实验结果

研究问题

  • RQ1ChatGPT能否在零样本提示下,基于既有的观察量规,准确可靠地对课堂教学进行评分?
  • RQ2与专家人类评分者相比,ChatGPT在课堂实录中识别教学亮点与改进机会的能力如何?
  • RQ3ChatGPT在提升学生推理方面的建议在多大程度上具有新颖性、可操作性,并忠实于实录内容?
  • RQ4ChatGPT的反馈建议中有多少比例描述了教师在实录中已实施的行动?
  • RQ5人类评分者对AI生成教练反馈质量的一致性水平如何?

主要发现

  • 即使提供了量规细节和推理提示,ChatGPT预测的评分与人工标注评分在所有CLASS和MQI条目上均表现出较低的相关性。
  • 50%至70%的ChatGPT对任务B(识别亮点与改进机会)的回应被专家评分者评为缺乏洞察力。
  • 任务B中35%至50%的回应被评定为与提示或实录内容无关。
  • 任务C中82%的建议描述了教师在实录中已实施的行动,表明存在高度冗余。
  • 人类评分者对模型输出质量的一致性仅为中等水平,相关性加权Cohen’s kappa为0.23,忠实度为0.36,洞察力为0.37。
  • 尽管标注者间一致性较低,但模型输出始终被评定为比有洞察力的反馈更相关、更具可操作性,表明在生成新颖、具有教学深度的反馈方面仍有改进空间。
((a))
((a))

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。