[논문 리뷰] Are Large Language Models a Threat to Digital Public Goods? Evidence from Activity on Stack Overflow
본 연구는 ChatGPT 출시 이후 Stack Overflow 게시가 대조 실험 플랫폼에 비해 약 16% 감소했고(6개월 사이에 ~25%로 상승), 게시 투표에는 유의한 변화가 없으며 더 인기 있는 언어에서 더 큰 감소가 나타났다는 것을 보여준다.
Large language models like ChatGPT efficiently provide users with information about various topics, presenting a potential substitute for searching the web and asking people for help online. But since users interact privately with the model, these models may drastically reduce the amount of publicly available human-generated data and knowledge resources. This substitution can present a significant problem in securing training data for future models. In this work, we investigate how the release of ChatGPT changed human-generated open data on the web by analyzing the activity on Stack Overflow, the leading online Q\&A platform for computer programming. We find that relative to its Russian and Chinese counterparts, where access to ChatGPT is limited, and to similar forums for mathematics, where ChatGPT is less capable, activity on Stack Overflow significantly decreased. A difference-in-differences model estimates a 16\% decrease in weekly posts on Stack Overflow. This effect increases in magnitude over time, and is larger for posts related to the most widely used programming languages. Posts made after ChatGPT get similar voting scores than before, suggesting that ChatGPT is not merely displacing duplicate or low-quality content. These results suggest that more users are adopting large language models to answer questions and they are better substitutes for Stack Overflow for languages for which they have more training data. Using models like ChatGPT may be more efficient for solving certain programming problems, but its widespread adoption and the resulting shift away from public exchange on the web will limit the open data people and models can learn from in the future.
연구 동기 및 목표
- ChatGPT와 같은 LLM이 Q&A 플랫폼에서 인간이 생성한 개방 데이터의 대체제가 되는지 평가한다.
- ChatGPT 출시 이후 차이의 차이(DID) 설계를 사용하여 Stack Overflow 게시 활동의 변화를 정량화한다.
- 투표 데이터를 통해 콘텐츠 품질에 미치는 변화를 분석한다.
- GitHub에서의 언어 인기와 개발자 급여 데이터와의 관계를 통해 프로그래밍 언어별 이질성을 탐구한다.
제안 방법
- Stack Overflow를 Math Stack Exchange, Math Overflow, Russian Stack Overflow, Segmentfault의 4개 대체 플랫폼과 비교하는 차이의 차이 모델을 사용한다.
- 효과를 백분율 변화로 해석하기 위해 IHS 변환으로 주간 게시물을 모델링하고 플랫폼 고정 효과, 주간 고정 효과, 플랫폼별 추세를 포함한다.
- 대상인 Stack Overflow와 게시 후 기간의 상호작용을 통해 ChatGPT 이후 효과를 추정하고, 주간 특성 상호작용으로 사전 추세를 검정한다.
- 69개의 언어 태그 주제에 걸친 언어 수준 이질성을 살펴보기 위한 이벤트 연구 설계를 보완적으로 적용한다.
- ChatGPT 출시 전후의 게시 품질의 대리 지표로서 투표 데이터(찬성/반대 투표)를 분석한다.
- 언어별 추정 효과를 GitHub 언어 인기 및 개발자 급여 데이터와 상관시킨다.

실험 결과
연구 질문
- RQ1ChatGPT의 출시가 비교 대상이자 덜 영향을 받은 플랫폼에 비해 Stack Overflow 게시 활동을 감소시키는가?
- RQ2투표 활동으로 측정된 콘텐츠의 품질이 감소하는가, 아니면 대체로 안정적인가?
- RQ3ChatGPT의 효과가 프로그래밍 언어마다 다르게 나타나며, 이러한 차이가 언어의 인기 또는 시장 신호와 관련이 있는가?
주요 결과
- Chat Overflow 게시 활동은 ChatGPT 출시 이후 약 15.6% 감소했고, 6개월 내에는 대략 25%까지 상승했다.
- 투표 활동(찬성/반대)은 안정적으로 유지되어 게시물 품질이 평균적으로 감소하지 않았음을 시사한다.
- 언어별로 이질적인 효과가 나타났다: 더 널리 사용되는 언어(Python, JavaScript 등)에서 게시 활동의 감소가 더 큰 경향을 보였다.
- GitHub 저장소가 더 많은 언어를 가진 경우 ChatGPT 출시 이후 Stack Overflow 게시에 더 큰 부정적 영향을 경험하는 경향이 나타났다.
- 대체 사양 및 하위표본(예: 질문만, 주중 게시물)에서도 결과의 강건성이 확인된다.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.