Skip to main content
QUICK REVIEW

[论文解读] Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

Julia Becker, Nate Rush|ArXiv.org|Jul 12, 2025
Big Data and Business Intelligence被引用 4
一句话总结

一项对16位经验丰富的开源开发者完成246个现实世界任务的随机对照试验显示,允许使用AI工具比不允许时完成时间慢19%,与预期的提速相反。

ABSTRACT

Despite widespread adoption, the impact of AI tools on software development in the wild remains understudied. We conduct a randomized controlled trial (RCT) to understand how AI tools at the February-June 2025 frontier affect the productivity of experienced open-source developers. 16 developers with moderate AI experience complete 246 tasks in mature projects on which they have an average of 5 years of prior experience. Each task is randomly assigned to allow or disallow usage of early 2025 AI tools. When AI tools are allowed, developers primarily use Cursor Pro, a popular code editor, and Claude 3.5/3.7 Sonnet. Before starting tasks, developers forecast that allowing AI will reduce completion time by 24%. After completing the study, developers estimate that allowing AI reduced completion time by 20%. Surprisingly, we find that allowing AI actually increases completion time by 19%--AI tooling slowed developers down. This slowdown also contradicts predictions from experts in economics (39% shorter) and ML (38% shorter). To understand this result, we collect and evaluate evidence for 20 properties of our setting that a priori could contribute to the observed slowdown effect--for example, the size and quality standards of projects, or prior developer experience with AI tooling. Although the influence of experimental artifacts cannot be entirely ruled out, the robustness of the slowdown effect across our analyses suggests it is unlikely to primarily be a function of our experimental design.

研究动机与目标

  • 评估早期到中期2025年的AI工具对经验丰富的开源开发者在现实世界中的真实生产力影响。
  • 使用固定的结果测量,比较允许AI与不允许AI时的任务完成时间。
  • 调查AI工具降低工作效率的原因,并识别导致因素的设定因素。

提出的方法

  • 对16位开发者在成熟的开源代码库上执行246个真实问题,进行随机对照试验。
  • 在随机分组前定义问题,并将其分配到AI允许或AI不允许的条件中。
  • 开发者使用Cursor Pro和Claude 3.5/3.7 Sonnet;结果基于总实现时间。
  • 在任务前后收集问题难度和AI影响的预测。
  • 丰富的数据源包括屏幕记录、代码库分析、访谈和调查。
  • 使用回归分析(对数线性)来估计总实现时间的百分比变化S。

实验结果

研究问题

  • RQ1允许早期到中期2025年的AI工具对经验丰富的开发者完成现实世界问题的总时间有何影响?
  • RQ2开发者与专家对AI影响的预期与观测结果有何异同?
  • RQ3在设定中哪些因素促成在允许AI时观察到的放慢?

主要发现

  • AI允许的任务平均完成时间比AI不允许的任务多19%。
  • 开发者预测AI会将时间缩短24%,但研究后他们估计比观察到的放慢多出20%的提速。
  • 专家(机器学习和经济学)预测的提速要远高于观察结果,分别为39%和38%。
  • 放慢的现象在多项分析中都显示出稳健性,尽管在21个候选因素中,各因素强度不同。
  • 有证据表明有5个因素促成放慢,对10个因素的证据混杂/不清楚,对6个因素的证据是否定。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。