Skip to main content
QUICK REVIEW

[论文解读] Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings

Veronika Hackl, Alexandra Elena Müller|arXiv (Cornell University)|Aug 3, 2023
Explainable Artificial Intelligence (XAI)参考文献 26被引用 4
一句话总结

本研究评估了GPT-4在高等教育宏观经济学任务中对学习者回答的一致性评分能力,采用结构化提示来评估内容与风格维度在多次评估中的表现。GPT-4表现出高度的评分者间信度(ICC 0.94–0.99),并有效区分了内容与风格,内容评分在风格重述下保持稳定,而风格评分在风格不充分时下降,表明其具备稳健且一致的评估能力。

ABSTRACT

This study investigates the consistency of feedback ratings generated by OpenAI's GPT-4, a state-of-the-art artificial intelligence language model, across multiple iterations, time spans and stylistic variations. The model rated responses to tasks within the Higher Education (HE) subject domain of macroeconomics in terms of their content and style. Statistical analysis was conducted in order to learn more about the interrater reliability, consistency of the ratings across iterations and the correlation between ratings in terms of content and style. The results revealed a high interrater reliability with ICC scores ranging between 0.94 and 0.99 for different timespans, suggesting that GPT-4 is capable of generating consistent ratings across repetitions with a clear prompt. Style and content ratings show a high correlation of 0.87. When applying a non-adequate style the average content ratings remained constant, while style ratings decreased, which indicates that the large language model (LLM) effectively distinguishes between these two criteria during evaluation. The prompt used in this study is furthermore presented and explained. Further research is necessary to assess the robustness and reliability of AI models in various use cases.

研究动机与目标

  • 评估GPT-4在高等教育背景下,对同一回答在多次迭代、不同时间跨度及风格变化下的评分一致性。
  • 评估GPT-4作为形成性评估评分者的可靠性,特别是对学生回答内容与风格维度的评分能力。
  • 检验GPT-4是否能可靠地区分书面回答中的内容质量与风格质量。
  • 为将AI生成的反馈整合到真实教育场景中(如BMBF资助的DeepWrite项目)提供基础支持。
  • 通过分析评分稳定性与相关性,为未来提示工程及AI在教育评估中的整合提供参考。

提出的方法

  • 使用标准化、详细的提示,要求GPT-4从内容与风格两个维度对宏观经济学学生回答进行评分。
  • 在多个时间点收集评分,并对同一回答的改写版本进行评估,以检验在变化条件下的评分一致性。
  • 计算组内相关系数(ICC)以衡量多次评估中评分者间的一致性。
  • 计算皮尔逊相关系数,以评估内容评分与风格评分之间的关系。
  • 分析评分分布的偏度,以检测GPT-4评分行为中是否存在偏差或倾向性。
  • 统计分析包括比较不同时间间隔与回答变体下的ICC值,以评估评分的稳定性。
Figure 1: Flowchart: Generated texts
Figure 1: Flowchart: Generated texts

实验结果

研究问题

  • RQ1GPT-4在长时间跨度内对同一回答的多次评估中,评分是否保持一致?
  • RQ2GPT-4在多大程度上能可靠地区分学生写作中的内容质量与风格质量?
  • RQ3回答风格的变化如何影响GPT-4对内容与风格的评分?
  • RQ4GPT-4评分的评分者间信度(以ICC衡量)在不同时间跨度与回答变体中如何表现?
  • RQ5评分分布的偏度是否表明GPT-4评分存在系统性偏差?

主要发现

  • GPT-4在不同时间跨度下表现出高度的评分者间信度,ICC分数在0.94至0.99之间,表明其在重复评分中具有极强的一致性。
  • 内容评分与风格评分之间的相关系数高达0.87,表明两者虽有关联,但被作为独立的评估标准处理。
  • 当回答被改写为非充分风格时,内容评分保持稳定,而风格评分下降,证实GPT-4能够有效区分内容与风格。
  • 内容评分呈现轻微右偏(平均偏度≈0.10),表明评分存在轻微宽松或平均分偏高的倾向。
  • 风格评分呈现左偏(平均偏度≈-0.03),表明其在风格评估中标准更严格或更保守。
  • 当评分在数周时间间隔后进行比较时,ICC值有所降低,表明长期间隔下一致性下降,但仍在高信度范围内。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。