Skip to main content
QUICK REVIEW

[논문 리뷰] Assessing the Ability of ChatGPT to Screen Articles for Systematic Reviews

Eugene Syriani, István Dávid|arXiv (Cornell University)|2023. 07. 12.
Scientific Computing and Data Management인용 수 17
한 줄 요약

이 논문은 체계적 고찰에서 논문 선별을 위한 ChatGPT의 일관성, 분류 성능 및 일반화 가능성을 평가하고, 이를 전통적 분류기와 비교하며 도구 통합 고려사항을 논의한다.

ABSTRACT

By organizing knowledge within a research field, Systematic Reviews (SR) provide valuable leads to steer research. Evidence suggests that SRs have become first-class artifacts in software engineering. However, the tedious manual effort associated with the screening phase of SRs renders these studies a costly and error-prone endeavor. While screening has traditionally been considered not amenable to automation, the advent of generative AI-driven chatbots, backed with large language models is set to disrupt the field. In this report, we propose an approach to leverage these novel technological developments for automating the screening of SRs. We assess the consistency, classification performance, and generalizability of ChatGPT in screening articles for SRs and compare these figures with those of traditional classifiers used in SR automation. Our results indicate that ChatGPT is a viable option to automate the SR processes, but requires careful considerations from developers when integrating ChatGPT into their SR tools.

연구 동기 및 목표

  • 소프트웨어 엔지니어링 분야에서 체계적 고찰의 선별 단계를 자동화할 필요성을 제기한다.
  • 선별 작업에서 ChatGPT의 일관성, 성능 및 일반화 가능성을 평가한다.
  • SR 자동화에 사용되는 전통 머신러닝 베이스라인과 ChatGPT를 비교한다.
  • 실증적 결과를 바탕으로 SR 도구에 ChatGPT를 통합하는 지침을 제시한다.

제안 방법

  • 선별 결정의 기준 진실로 ReLiS 기반 SR 코퍼스를 사용한다.
  • 단어2벡(features)을 활용한 기사 제목/초록으로 학습된 전통 분류기(LR, RF, CNB, SVC)로 기준선을 설정한다.
  • ChatGPT용 프롬프트를 설계하고 그의 선별 판단을 기준 진실과 대조 평가한다.
  • 불균형 데이터 성능 지표(MCC, F2, 균형 정확도 등)를 사용해 ChatGPT 결과를 기준선과 비교한다.
  • 여러 실행에 걸친 Fleiss’ Kappa로 일관성을 평가하여 ChatGPT 판단의 안정성을 측정한다.

실험 결과

연구 질문

  • RQ1RQ1: 특정 논문에 대한 ChatGPT의 선별 판단은 실행 간에 얼마나 일관성 있는가?
  • RQ2RQ2: ChatGPT의 분류 성능은 전통적 SR 자동화 분류기와 어떻게 비교되는가?
  • RQ3RQ3: ChatGPT의 선별 판단은 다양한 소프트웨어 공학 SR 데이터셋에서 얼마나 일반화되는가?

주요 결과

  • ChatGPT는 추가 훈련 없이도 SR 선별에서 전통적 머신러닝 방법의 성능에 필적할 수 있다.
  • LLMs like ChatGPT show potential to revolutionize SR automation, but require careful integration considerations for tool developers.
  • 다양한 SR 데이터셋에서 포함/제외 비율과 충돌이 다르게 나타나 일반화를 테스트했고, 견고한 프롬프트 엔지니어링의 필요성을 강조했다.
  • 본 연구는 ReLiS 프로젝트의 기준 진실 데이터를 활용해 포함/제외 결정의 신뢰할 수 있는 평가를 보장한다.
  • 프롬프트와 하이퍼파라미터(온도, 토큰 한도)는 일관되고 최소한의 응답 결과(포함/제외)에 결정적이다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.