[Paper Review] Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
An RCT of 16 experienced open-source developers completing 246 real-world tasks shows that allowing AI tooling slows completion time by 19%, contrary to expected speedups.
Despite widespread adoption, the impact of AI tools on software development in the wild remains understudied. We conduct a randomized controlled trial (RCT) to understand how AI tools at the February-June 2025 frontier affect the productivity of experienced open-source developers. 16 developers with moderate AI experience complete 246 tasks in mature projects on which they have an average of 5 years of prior experience. Each task is randomly assigned to allow or disallow usage of early 2025 AI tools. When AI tools are allowed, developers primarily use Cursor Pro, a popular code editor, and Claude 3.5/3.7 Sonnet. Before starting tasks, developers forecast that allowing AI will reduce completion time by 24%. After completing the study, developers estimate that allowing AI reduced completion time by 20%. Surprisingly, we find that allowing AI actually increases completion time by 19%--AI tooling slowed developers down. This slowdown also contradicts predictions from experts in economics (39% shorter) and ML (38% shorter). To understand this result, we collect and evaluate evidence for 20 properties of our setting that a priori could contribute to the observed slowdown effect--for example, the size and quality standards of projects, or prior developer experience with AI tooling. Although the influence of experimental artifacts cannot be entirely ruled out, the robustness of the slowdown effect across our analyses suggests it is unlikely to primarily be a function of our experimental design.
Motivation & Objective
- Assess the real-world productivity impact of early-2025 AI tools on experienced open-source developers.
- Measure task completion time with AI allowed vs. AI disallowed using fixed outcome measures.
- Investigate why AI tooling slowed work and identify contributing setting factors.
Proposed method
- Randomized controlled trial with 16 developers performing 246 real issues on mature open-source repositories.
- Issues defined prior to randomization and assigned to AI-allowed or AI-disallowed conditions.
- Developers used Cursor Pro and Claude 3.5/3.7 Sonnet; outcomes are based on total implementation time.
- Forecasts of issue difficulty and AI impact were collected before and after tasks.
- Rich data sources include screen recordings, repository analytics, interviews, and surveys.
- Regression analysis (log-linear) used to estimate percent change in total implementation time S.
Experimental results
Research questions
- RQ1What is the effect of allowing early-2025 AI tooling on the total time to complete real-world issues for experienced developers?
- RQ2How do developer and expert expectations of AI impact compare with observed outcomes?
- RQ3Which factors in the setting contribute to any observed slowdown when AI is allowed?
Key findings
- AI-allowed tasks took 19% longer on average to complete than AI-disallowed tasks.
- Developers forecast AI would shorten time by 24%, but post-study, they estimated a 20% speedup versus observed slowdown.
- Experts (ML and economics) forecast substantially larger speedups (39% and 38%, respectively) than observed.
- The slowdown appears robust across multiple analyses, though factors vary in strength across the 21 candidate factors considered.
- There is evidence that 5 factors contribute to slowdown, with mixed/unclear evidence for 10 factors and evidence against 6 factors.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.