Skip to main content
QUICK REVIEW

[논문 리뷰] Exploring the use of a Large Language Model for data extraction in systematic reviews: a rapid feasibility study

Lena Schmidt, Kaitlyn Hair|arXiv (Cornell University)|2024. 05. 23.
Artificial Intelligence in Healthcare인용 수 9
한 줄 요약

논문은 데이터 추출 자동화를 위한 빠른 현황 연구를 GPT-4를 사용하여 체계적 문헌고찰에 대해 수행하고, 영역 간 정확도와 변동성을 평가하며 도전과제와 평가 방법을 강조한다.

ABSTRACT

This paper describes a rapid feasibility study of using GPT-4, a large language model (LLM), to (semi)automate data extraction in systematic reviews. Despite the recent surge of interest in LLMs there is still a lack of understanding of how to design LLM-based automation tools and how to robustly evaluate their performance. During the 2023 Evidence Synthesis Hackathon we conducted two feasibility studies. Firstly, to automatically extract study characteristics from human clinical, animal, and social science domain studies. We used two studies from each category for prompt-development; and ten for evaluation. Secondly, we used the LLM to predict Participants, Interventions, Controls and Outcomes (PICOs) labelled within 100 abstracts in the EBM-NLP dataset. Overall, results indicated an accuracy of around 80%, with some variability between domains (82% for human clinical, 80% for animal, and 72% for studies of human social sciences). Causal inference methods and study design were the data extraction items with the most errors. In the PICO study, participants and intervention/control showed high accuracy (>80%), outcomes were more challenging. Evaluation was done manually; scoring methods such as BLEU and ROUGE showed limited value. We observed variability in the LLMs predictions and changes in response quality. This paper presents a template for future evaluations of LLMs in the context of data extraction for systematic review automation. Our results show that there might be value in using LLMs, for example as second or third reviewers. However, caution is advised when integrating models such as GPT-4 into tools. Further research on stability and reliability in practical settings is warranted for each type of data that is processed by the LLM.

연구 동기 및 목표

  • 대형언어모형(LLMs)이 체계적 고찰에서 데이터 추출을 돕는 방식에 대한 동기 부여와 탐색.
  • LLM 기반 추출을 위한 프롬프트 템플릿과 평가 프로토콜을 개발.
  • 교차 도메인 성능 평가(임상 인간, 동물, 그리고 인간 사회과학).
  • LLMs가 잘 작동하는 영역과 오류가 더 자주 발생하는 영역을 식별하여 향후 도구 설계에 반영.

제안 방법

  • 2023 Evidence Synthesis Hackathon 동안 수행된 두 건의 현황 타당성 연구.
  • 첫 번째 연구: 도메인 연구로부터 연구 특성의 자동 추출; 프롬프트 개발에 도메인당 두 건의 연구를 사용하고 평가에 열 건을 사용.
  • 두 번째 연구: EBM-NLP 데이터세트의 100개 초록에서 PICO(참여자, 중재, 대조, 결과)를 LLM이 예측.
  • 평가는 BLEU/ROUGE 지표에 의존하기보다는 수동 평가였다.
  • 예측의 변동성과 응답 품질의 변화 식별.
  • 데이터 추출 맥락에서 LLM를 평가하기 위한 템플릿 제시.

실험 결과

연구 질문

  • RQ1GPT-4가 임상, 동물, 그리고 사회과학 연구에서 연구 특성을 정확하게 추출할 수 있는가?
  • RQ2초록에서 PICO를 식별하는 데 LLM의 성능은 어느 정도이며 어떤 구성요소가 가장 오류가 많이 발생하는가?
  • RQ3BLEU/ROUGE를 넘어 LLM 기반 데이터 추출에 적합한 평가 방법은 무엇인가?
  • RQ4데이터 추출 워크플로에 LLM를 통합할 때의 안정성 및 신뢰성 고려사항은 무엇인가?

주요 결과

  • 데이터 추출 작업의 전반적 정확도 약 80% 수준이며 영역 간 변동성 있음(임상 인간 82%, 동물 80%, 사회과학 72%).
  • 인과 추론 방법과 연구 설계가 가장 많은 오류를 보인 데이터 추출 항목이었다.
  • PICO 연구에서 참여자와 중재/대조는 높은 정확도(>80%)를 보였으나 결과는 더 어려웠다.
  • 평가는 수동으로 이루어졌고 BLEU 및 ROUGE와 같은 전통적 지표는 제한적 가치를 보였다.
  • 예측의 변동성과 실행 간 응답 품질의 변화가 있었다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.