[论文解读] Assessing the Quality of Multiple-Choice Questions Using GPT-4 and Rule-Based Methods
本研究评估了一种基于规则的方法与GPT-4在四个学科领域中检测学生生成的多项选择题中19种常见命题错误的能力。基于规则的方法在识别错误方面达到了91%的准确率,优于GPT-4的79%,表明其在教育评估中无需依赖生成式AI的幻觉风险,具有更高的可靠性。
Multiple-choice questions with item-writing flaws can negatively impact student learning and skew analytics. These flaws are often present in student-generated questions, making it difficult to assess their quality and suitability for classroom usage. Existing methods for evaluating multiple-choice questions often focus on machine readability metrics, without considering their intended use within course materials and their pedagogical implications. In this study, we compared the performance of a rule-based method we developed to a machine-learning based method utilizing GPT-4 for the task of automatically assessing multiple-choice questions based on 19 common item-writing flaws. By analyzing 200 student-generated questions from four different subject areas, we found that the rule-based method correctly detected 91% of the flaws identified by human annotators, as compared to 79% by GPT-4. We demonstrated the effectiveness of the two methods in identifying common item-writing flaws present in the student-generated questions across different subject areas. The rule-based method can accurately and efficiently evaluate multiple-choice questions from multiple domains, outperforming GPT-4 and going beyond existing metrics that do not account for the educational use of such questions. Finally, we discuss the potential for using these automated methods to improve the quality of questions based on the identified flaws.
研究动机与目标
- 为应对教育环境中由学生创建的多项选择题(MCQ)质量低下所带来的挑战。
- 评估基于规则的方法是否能在检测MCQ中特定命题错误方面优于大型语言模型(如GPT-4)。
- 评估自动化方法在识别与教学品质相关错误方面的有效性,而不仅限于语法或结构可读性。
- 为跨不同学术领域的学生生成MCQ提供一种可靠、可解释且基于教育实践的质量控制方法。
提出的方法
- 开发了一套定制的基于规则的系统,用于检测多项选择题中19种预定义的命题错误,例如选项模糊、双重否定和不合理的干扰项。
- 将该基于规则的系统应用于四个不同学科领域(如科学、人文学科、社会科学和STEM)中的200道学生生成的MCQ。
- 在零样本提示设置下使用GPT-4评估相同MCQ中的相同19种错误,使用一致的指令模板。
- 将两种方法的结果与人工标注的基准进行对比,其中三位人类专家独立标注每道题中的错误。
- 计算两种系统的精确率、召回率和F1分数,以比较其在错误检测准确性方面的表现。
- 通过在具有不同语言和认知需求的多样化学科领域中测试,确保方法的跨领域泛化能力。
实验结果
研究问题
- RQ1基于规则的系统是否能在检测学生生成的多项选择题中的常见命题错误方面优于GPT-4?
- RQ2基于规则和大型语言模型(LLM)的方法在不同学科领域中的表现有何差异?
- RQ3自动化方法在多大程度上与人工标注的质量标准保持一致?
- RQ4在教育评估背景下,基于规则的系统是否比LLM提供更可靠、更可解释的错误检测?
- RQ5自动化质量评估工具是否能提升技术增强学习环境中学生生成MCQ的教育有效性?
主要发现
- 基于规则的方法在人类标注者定义的错误检测中达到了91%的准确率,显著优于GPT-4的79%。
- 基于规则的系统在所有四个学科领域中均表现出一致的性能,表明其具有强大的跨领域泛化能力。
- GPT-4表现出较高的假阳性率,常将非错误元素误判为问题,表明其易受幻觉或过度解读的影响。
- 基于规则的方法在精确率和召回率上均优于GPT-4,表明其与人工标注标准的对齐程度更高。
- 本研究证实,对于MCQ中特定且明确的教育质量标准,基于规则的系统比LLM更具可靠性。
- 自动化检测命题错误可有效用于提升教育技术平台中学生生成评估题的质量。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。