Skip to main content
QUICK REVIEW

[논문 리뷰] Bio-SIEVE: Exploring Instruction Tuning Large Language Models for Systematic Review Automation

Ambrose Robinson, William Thorne|arXiv (Cornell University)|2023. 08. 12.
Meta-analysis and systematic reviews인용 수 7
한 줄 요약

Bio-SIEVE는 생물의학적 체계적 문헌 고찰에서 자동 문헌 선별을 위한 지시 조정된 LLM을 도입하여, 세부적인 선별 기준을 사용해 연구의 관련성을 정확하게 분류함으로써 ChatGPT와 전통적 모델을 능가한다. 이는 다양한 의료 분야에서 뛰어난 일반화 성능을 보이며 배제 이유를 설명할 수 있도록 하여 재학습 없이도 영향력 있고 영향력 있는 영역에 맞는 제로샷 솔루션을 제공한다.

ABSTRACT

Medical systematic reviews can be very costly and resource intensive. We explore how Large Language Models (LLMs) can support and be trained to perform literature screening when provided with a detailed set of selection criteria. Specifically, we instruction tune LLaMA and Guanaco models to perform abstract screening for medical systematic reviews. Our best model, Bio-SIEVE, outperforms both ChatGPT and trained traditional approaches, and generalises better across medical domains. However, there remains the challenge of adapting the model to safety-first scenarios. We also explore the impact of multi-task training with Bio-SIEVE-Multi, including tasks such as PICO extraction and exclusion reasoning, but find that it is unable to match single-task Bio-SIEVE's performance. We see Bio-SIEVE as an important step towards specialising LLMs for the biomedical systematic review process and explore its future developmental opportunities. We release our models, code and a list of DOIs to reconstruct our dataset for reproducibility.

연구 동기 및 목표

  • 생물의학적 체계적 문헌 고찰에서 수작업 문헌 선별의 높은 비용과 시간 부담을 해결하기 위해.
  • 재학습 없이 다양한 의료 분야로 일반화되는 제로샷 지시 조정 LLM을 개발하기 위해.
  • 복잡한 다중 기준 선별 규칙을 사용해 포함/배제 분류를 정확히 수행하기 위해.
  • 투명성 향상과 리뷰어의 부담 감소를 위해 배제 이유 분석을 새로운 과제로 탐색하기 위해.
  • 이전 지식 전이 효과를 확보하기 위해 PICO 추출 및 배제 이유 설명을 포함한 다중 과제 학습 평가하기 위해.

제안 방법

  • Cochrane 리뷰에서 유래한 정제된 데이터셋을 기반으로 지시 조정을 통해 미세조정된 LLaMA 및 Guanaco 모델을 사용함.
  • 특정 리뷰 기준에 따라 포함/배제를 레이블링한 摘要 데이터셋을 구축함. 이는 배제 이유 포함.
  • 선별 및 설명 과제의 모델 행동을 이끌기 위해 상세한 자연어 프롬프트를 사용한 지시 조정을 적용함.
  • 교차 과제 지식 전이를 활용하기 위해 PICO 추출 및 배제 이유 설명을 포함한 다중 과제 학습을 탐색함.
  • 도메인 특화 선별 과제에 대해 대규모 모델을 효율적으로 미세조정하기 위해 LoRA(Low-Rank Adaptation)를 사용함.
  • 표준 NLP 메트릭(예: F1, 정밀도, 재현율, 인간 선호도 순위)을 사용해 성능 평가함.
Figure 1: A simple representation of the systematic review process depicting the stage which Bio-SIEVE aims to assist. The black funnels are the monotonous and highly resource intensive bottlenecks of the process.
Figure 1: A simple representation of the systematic review process depicting the stage which Bio-SIEVE aims to assist. The black funnels are the monotonous and highly resource intensive bottlenecks of the process.

실험 결과

연구 질문

  • RQ1지시 조정된 LLM은 ChatGPT와 같은 제로샷 LLM보다 체계적 문헌 고찰의 개요 선별에서 뛰어난 성능을 보일 수 있는가?
  • RQ2도메인 특화 기준에 대한 지시 조정이 다양한 생물의학 주제 간 일반화 능력을 향상시키는가?
  • RQ3배제 이유 분석을 생성 과제로 효과적으로 모델링하여 선별 결정의 투명성을 지원할 수 있는가?
  • RQ4PICO 추출 및 배제 이유 설명을 포함한 다중 과제 학습이 전체 선별 성능 향상에 기여하는가?
  • RQ5민감도 및 특이도 측면에서 Bio-SIEVE는 전통적 기계학습 모델보다 어떻게 비교되는가?

주요 결과

  • Bio-SIEVE는 포함/배제 분류에서 ChatGPT와 SVM과 같은 전통적 모델을 모두 능가하여 더 높은 F1 및 재현율을 달성함.
  • 재학습 없이도 다양한 의료 분야로 효과적으로 일반화되어 강력한 제로샷 일반화 능력을 보임.
  • 배제 이유 분석이 성공적으로 모델링되어 연구를 기각한 타당하고 맥락에 부합하는 이유를 생성할 수 있음.
  • 다중 과제 학습(Bio-SIEVE-Multi)은 전이 효과가 제한적이며, 선별 과제에서는 단일 과제 Bio-SIEVE보다 성능이 열 劣함.
  • Bio-SIEVE-Multi는 포함 이유 분석에서는 잠재력을 보였지만 인간 선호도 순위에서 ChatGPT의 성능을 따라잡지 못함.
  • 구두로 관련성이 없는 사례(예: 구강 건강 고찰에서 근육 외상 연구 제외)를 정밀하게 걸러내는 데 높은 정확도를 달성함.
Figure 2: The topic distribution of the inclusion/exclusion classification samples in the train and test splits of the Instruct Cochrane dataset.
Figure 2: The topic distribution of the inclusion/exclusion classification samples in the train and test splits of the Instruct Cochrane dataset.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.