[Paper Review] AI Meets the Classroom: When Do Large Language Models Harm Learning?
The paper investigates how access to large language models (LLMs) like ChatGPT affects learning in coding education, finding that LLMs can boost learning when used for explanations but harm it when used to provide solutions, with stronger negative effects for students with less prior knowledge.
The effect of large language models (LLMs) in education is debated: Previous research shows that LLMs can help as well as hurt learning. In two pre-registered and incentivized laboratory experiments, we find no effect of LLMs on overall learning outcomes. In exploratory analyses and a field study, we provide evidence that the effect of LLMs on learning outcomes depends on usage behavior. Students who substitute some of their learning activities with LLMs (e.g., by generating solutions to exercises) increase the volume of topics they can learn about but decrease their understanding of each topic. Students who complement their learning activities with LLMs (e.g., by asking for explanations) do not increase topic volume but do increase their understanding. We also observe that LLMs widen the gap between students with low and high prior knowledge. While LLMs show great potential to improve learning, their use must be tailored to the educational context and students' needs.
Motivation & Objective
- Explore how LLM access impacts learning outcomes in coding education across field and laboratory settings.
- Identify mechanisms: explanation-based use versus solution-seeking and the role of copy-paste functionality.
- Assess heterogeneity by student ability and prior knowledge.
- Examine perceived vs. actual learning progress under LLM usage.
- Provide policy-relevant guidance on leveraging LLMs as learning supports while mitigating pitfalls.
Proposed method
- Three-study design combining observational field data and two incentivized, pre-registered laboratory experiments.
- Field data analysis using two university programming courses with a two-way fixed effects model and instrumental variables.
- Measurement of LLM usage via ChatGPT similarity between student code and ChatGPT-generated code as a proxy for usage.
- Experimental manipulation of LLM access and a copy-paste feature to test causal effects and mechanisms.
- Pre-registration and use of Outages-based IVs to identify exogenous variation in LLM usage.
- Use of Python programming tasks with standardized pre-, learning-, and post-tests to measure learning progress.

Experimental results
Research questions
- RQ1Does access to LLMs improve learning outcomes in coding when used as a tutor or explainer?
- RQ2Does LLM usage that facilitates solution-seeking impair subsequent learning?
- RQ3How does prior coding knowledge interact with the effects of LLM usage?
- RQ4To what extent do students overestimate their learning progress when using LLMs?
Key findings
- LLM-generated explanations improve learning, while solving exercises with LLMs can harm subsequent learning.
- In field data, a current-question ChatGPT solution raises a grade, while cumulative ChatGPT similarity adversely affects later performance.
- Instrumental variable analysis confirms a negative effect of cumulative ChatGPT usage on learning, suggesting robust negative learning effects for overreliance.
- Weaker students benefit more from LLM access, consistent with prior ability heterogeneity findings.
- Participants report higher perceived learning progress than actual progress, indicating overconfidence in LLM-assisted learning.
- Study 3 (not fully shown) further isolates mechanisms and supports the potential for LLMs as effective learning aids when used appropriately.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.