Skip to main content
QUICK REVIEW

[논문 리뷰] Natural Language Programming in Medicine: Administering Evidence Based Clinical Workflows with Autonomous Agents Powered by Generative Large Language Models

Akhil Vaid, Joshua Lampert|arXiv (Cornell University)|2024. 01. 05.
Artificial Intelligence in Healthcare and Education인용 수 7
한 줄 요약

이 논문은 생성형 LLM으로 구동되는 자율 에이전트를 사용해 시뮬레이션된 3차 진료 환경에서 근거 기반 임상 워크플로를 관리하고, 독점 모델과 오픈 소스 모델을 RAG와 비교하며 인간 감독의 필요성과 NLP 기반 행동 수정의 필요성을 강조합니다.

ABSTRACT

Generative Large Language Models (LLMs) hold significant promise in healthcare, demonstrating capabilities such as passing medical licensing exams and providing clinical knowledge. However, their current use as information retrieval tools is limited by challenges like data staleness, resource demands, and occasional generation of incorrect information. This study assessed the potential of LLMs to function as autonomous agents in a simulated tertiary care medical center, using real-world clinical cases across multiple specialties. Both proprietary and open-source LLMs were evaluated, with Retrieval Augmented Generation (RAG) enhancing contextual relevance. Proprietary models, particularly GPT-4, generally outperformed open-source models, showing improved guideline adherence and more accurate responses with RAG. The manual evaluation by expert clinicians was crucial in validating models' outputs, underscoring the importance of human oversight in LLM operation. Further, the study emphasizes Natural Language Programming (NLP) as the appropriate paradigm for modifying model behavior, allowing for precise adjustments through tailored prompts and real-world interactions. This approach highlights the potential of LLMs to significantly enhance and supplement clinical decision-making, while also emphasizing the value of continuous expert involvement and the flexibility of NLP to ensure their reliability and effectiveness in healthcare settings.

연구 동기 및 목표

  • 의학에서 근거 기반 임상 워크플로를 실행하기 위한 자율 LLM 에이전트의 사용을 동기 부여하고 평가합니다.
  • 3차 진료 시뮬레이션에서 가이드라인 준수 및 응답 정확도 측면에서 독점 모델과 오픈 소스 LLM을 비교합니다.
  • 맥락 적합성과 의사결정 품질에 대한 Retrieval Augmented Generation (RAG)의 영향을 평가합니다.
  • 클리닉 맥락에서 모델 행동을 안전하게 조정하기 위한 실용적 패러다임으로서 Natural Language Programming(NLP)을 보여줍니다.

제안 방법

  • 여러 전문 분야에 걸친 실제 임상 사례를 이용한 3차 진료 메디컬 센터 시뮬레이션.
  • 자율 임상 업무 수행을 위해 독점 및 오픈 소스 LLM을 평가합니다.
  • 출력의 맥락 관련성을 높이기 위해 Retrieval Augmented Generation (RAG)을 도입합니다.
  • 모델 출력의 타당성을 검증하기 위해 전문가 의사 평가를 적용합니다.
  • prompt 및 실제 상호 작용을 통해 모델 행동을 수정하는 패러다임으로서 NLP를 옹호합니다.

실험 결과

연구 질문

  • RQ1시뮬레이션 병원 환경에서 여러 전문 분야에 걸쳐 자율 LLM 에이전트가 임상 가이드라인을 신뢰성 있게 준수할 수 있는가?
  • RQ2RAG를 사용할 때 GPT-4와 같은 독점 모델이 가이드라인 준수 및 정확도 측면에서 오픈 소스 모델보다 우수한가?
  • RQ3Retrieval Augmented Generation이 LLM 주도 임상 워크플로의 맥락 관련성 및 정확성을 개선하는가?
  • RQ4자율 의료 에이전트를 검증하고 감독하는 인간 전문가의 역할은 무엇인가?
  • RQ5자율 임상 에이전트를 신뢰성과 안전성 있게 조정하는 효과적인 방법으로서 Natural Language Programming이 가능한가?

주요 결과

  • 독점 모델, 특히 GPT-4는 RAG를 사용할 때 오픈 소스 모델에 비해 일반적으로 가이드라인 준수와 정확성에서 우수한 편입니다.
  • RAG는 의학 자율성 설정에서 응답의 맥락 관련성을 향상시킵니다.
  • 모델 출력의 타당성 확인 및 안전한 작동 보장을 위해 수동 전문가 의사 평가가 매우 중요합니다.
  • NL 프로그래밍은 맞춤형 프롬프트와 상호 작용을 통해 모델 행동을 정확하게 조정할 수 있게 합니다.
  • 이 접근 방식은 LLMS가 임상 의사결정을 보완할 잠재력을 보이지만 지속적인 전문가 참여가 필요함을 시사합니다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.