Skip to main content
QUICK REVIEW

[논문 리뷰] A Preliminary Study of o1 in Medicine: Are We Closer to an AI Doctor?

Yunfei Xie, Jianzhen Wu|arXiv (Cornell University)|2024. 09. 23.
Artificial Intelligence in Healthcare and Education인용 수 14
한 줄 요약

본 논문은 이해, 추론, 다국어 능력에 걸쳐 OpenAI의 o1 모델을 37개의 의학 데이터셋으로 평가하고, 의학적 이해 및 추론이 향상되었음을 보여주지만, 환각, 다국어 도전과제, 그리고 지표 불일치를 강조한다.

ABSTRACT

Large language models (LLMs) have exhibited remarkable capabilities across various domains and tasks, pushing the boundaries of our knowledge in learning and cognition. The latest model, OpenAI's o1, stands out as the first LLM with an internalized chain-of-thought technique using reinforcement learning strategies. While it has demonstrated surprisingly strong capabilities on various general language tasks, its performance in specialized fields such as medicine remains unknown. To this end, this report provides a comprehensive exploration of o1 on different medical scenarios, examining 3 key aspects: understanding, reasoning, and multilinguality. Specifically, our evaluation encompasses 6 tasks using data from 37 medical datasets, including two newly constructed and more challenging question-answering (QA) tasks based on professional medical quizzes from the New England Journal of Medicine (NEJM) and The Lancet. These datasets offer greater clinical relevance compared to standard medical QA benchmarks such as MedQA, translating more effectively into real-world clinical utility. Our analysis of o1 suggests that the enhanced reasoning ability of LLMs may (significantly) benefit their capability to understand various medical instructions and reason through complex clinical scenarios. Notably, o1 surpasses the previous GPT-4 in accuracy by an average of 6.2% and 6.6% across 19 datasets and two newly created complex QA scenarios. But meanwhile, we identify several weaknesses in both the model capability and the existing evaluation protocols, including hallucination, inconsistent multilingual ability, and discrepant metrics for evaluation. We release our raw data and model outputs at https://ucsc-vlaa.github.io/o1_medicine/ for future research.

연구 동기 및 목표

  • o1 모델의 향상된 추론 능력이 의학 영역으로 이전되는지 평가한다.
  • 다양한 데이터세트를 사용하여 의학적 이해, 추론 및 다국어 역량에서 o1를 평가한다.
  • 여러 의학 과제에서 o1를 GPT-4, GPT-3.5 및 오픈 소스 기반 기준들과 비교한다.
  • 향후 임상 인공지능 개발을 위한 모델 성능의 약점과 현재 평가 프로토콜의 한계를 식별한다.

제안 방법

  • 의학의 세 가지 측면에 걸친 37개 데이터셋(기존 35개 + 신규 2개)을 망라한 광범위한 평가 세트를 구성한다.
  • 직접 프롬팅, 사고 사슬(CoT), 소수 샷 프롬팅의 세 가지 프롬팅 전략을 사용한다; o1의 내부 CoT 학습을 고려할 때 CoT의 영향을 평가한다.
  • 여섯 가지 과제와 세 가지 측면에 대해 o1를 GPT-4, GPT-3.5, MEDITRON-70B, Llama3-8B와 비교한다.
  • 정확도(Accuracy), F1, BLEU, ROUGE, AlignScore, Mauve 등의 지표를 사용하여 다양한 과제 유형(이해, 추론, 다국어성)을 평가한다.
  • CoT, Self-Consistency, Reflex 등의 추가 프롬프트로 결과를 분석하여 프롬프트의 효과를 연구한다.

실험 결과

연구 질문

  • RQ1이전 모델과 비교했을 때 o1의 내부 사고 사슬과 강화 학습 훈련이 임상 이해 및 추론을 향상시키는가?
  • RQ2GPT-4, GPT-3.5, 및 오픈 소스 기준과 비교하여 o1은 의학적 이해, 추론, 다국어 과제에서 어떻게 작동하는가?
  • RQ3의학 맥락에서의 환각 및 다국어 도전과제를 포함한 o1의 한계는 무엇이며, 평가 지표가 모델 순위에 어떤 영향을 미치는가?

주요 결과

  • 다수의 데이터셋에서 이해 및 특정 추론 과제에 대해 o1이 일반적으로 GPT-4 및 GPT-3.5를 능가한다.
  • 새로운 NEJMQA 및 LancetQA 과제에서 o1은 GPT-4 및 GPT-3.5에 비해 명확한 정확도 향상을 보인다.
  • o1은 자유 형식 생성 과제에서 ROUGE-1 점수가 더 높고 요약 품질이 향상됨을 보여준다.
  • 환각은 여전히 o1의 도전 과제이며 다국어의 복잡한 시나리오는 다국어 추론의 격차를 드러낸다.
  • 평가 지표는 모델 간 불일치하는 순위를 낳아 강건하고 도메인 특화된 지표의 필요성을 부각시킨다.
  • CoT 프롬핑은 o1의 의학 지식 과제를 개선할 수 있지만 모든 과제 유형에 보편적으로 적용되지는 않는다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.