Skip to main content
QUICK REVIEW

[论文解读] AI Meets the Classroom: When Do Large Language Models Harm Learning?

Matthias Lehmann, Philipp B. Cornelius|arXiv (Cornell University)|Aug 29, 2024
Artificial Intelligence in Healthcare and Education被引用 5
一句话总结

本论文调查像 ChatGPT 这样的 大型语言模型(LLMs) 如何影响编码教育中的学习,发现当用于解释时 LLMs 能提升学习,而用于提供解答时会削弱学习,对先前知识较少的学生的负面效应更强。

ABSTRACT

The effect of large language models (LLMs) in education is debated: Previous research shows that LLMs can help as well as hurt learning. In two pre-registered and incentivized laboratory experiments, we find no effect of LLMs on overall learning outcomes. In exploratory analyses and a field study, we provide evidence that the effect of LLMs on learning outcomes depends on usage behavior. Students who substitute some of their learning activities with LLMs (e.g., by generating solutions to exercises) increase the volume of topics they can learn about but decrease their understanding of each topic. Students who complement their learning activities with LLMs (e.g., by asking for explanations) do not increase topic volume but do increase their understanding. We also observe that LLMs widen the gap between students with low and high prior knowledge. While LLMs show great potential to improve learning, their use must be tailored to the educational context and students' needs.

研究动机与目标

  • 在现场和实验室环境中探索 LLM 访问如何影响编程教育中的学习结果。
  • 识别机制:基于解释的使用 versus 寻求解答,以及复制粘贴功能的作用。
  • 评估学生能力和先前知识水平的异质性。
  • 在使用 LLM 的条件下考察感知学习进展与实际学习进展之间的差异。
  • 提供关于在学习支持中有效利用 LLM 同时减小风险的政策相关指导。

提出的方法

  • 三项研究设计,结合观察性现场数据与两项带激励、事前注册的实验室实验。
  • 使用两门大学编程课程的现场数据分析,采用双向固定效应模型和工具变量。
  • 通过学生代码与 ChatGPT 生成代码之间的相似度来衡量 LLM 的使用,作为使用的代理。
  • 对 LLM 访问和复制粘贴功能进行实验性操控,以检验因果效应与机制。
  • 事前注册并使用基于 outages 的工具变量来识别 LLM 使用中的外生变异。
  • 使用带有标准化前测、学习期和后测的 Python 编程任务来衡量学习进展。
AI Meets the Classroom: When Do Large Language Models Harm Learning?

实验结果

研究问题

  • RQ1在作为导师或解释者使用时,LLM 的访问是否能改善编程中的学习结果?
  • RQ2促使寻求解答的 LLM 使用是否会损害后续学习?
  • RQ3先前的编码知识如何与 LLM 使用效果相互作用?
  • RQ4在使用 LLM 时,学生在多大程度上高估了自己的学习进展?

主要发现

  • LLM 生成的解释能提升学习,而用 LLMs 解决练习则可能损害后续学习。
  • 在现场数据中,当前问题的 ChatGPT 解答提高了分数,而累计的 ChatGPT 相似度会对后期表现产生不利影响。
  • 工具变量分析证实累计使用 ChatGPT 对学习的负面影响,提示过度依赖的负面学习效应更为显著。
  • 学习能力较弱的学生从 LLM 访问中获益更多,与先前的能力异质性发现一致。
  • 参与者报告的感知学习进展高于实际进展,表明对 LLM 辅助学习的过度自信。
  • 研究3(未完全展示)进一步分离机制,并支持在适当使用时将 LLM 作为有效学习辅助工具的潜力。
AI Meets the Classroom: When Do Large Language Models Harm Learning?

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。