Skip to main content
QUICK REVIEW

[논문 리뷰] Polaris: A Safety-focused LLM Constellation Architecture for Healthcare

Subhabrata Mukherjee, Paul Gamble|arXiv (Cornell University)|2024. 03. 20.
Quality and Safety in Healthcare인용 수 19
한 줄 요약

Polaris는 안전 중심의 다중 에이전트 LLM 구성으로 실시간, 환자 대면 의료 대화를 제공하며, 주요 에이전트와 전문 지원 에이전트가 의료 안전성을 강화하고 환각을 줄인다.

ABSTRACT

We develop Polaris, the first safety-focused LLM constellation for real-time patient-AI healthcare conversations. Unlike prior LLM works in healthcare focusing on tasks like question answering, our work specifically focuses on long multi-turn voice conversations. Our one-trillion parameter constellation system is composed of several multibillion parameter LLMs as co-operative agents: a stateful primary agent that focuses on driving an engaging conversation and several specialist support agents focused on healthcare tasks performed by nurses to increase safety and reduce hallucinations. We develop a sophisticated training protocol for iterative co-training of the agents that optimize for diverse objectives. We train our models on proprietary data, clinical care plans, healthcare regulatory documents, medical manuals, and other medical reasoning documents. We align our models to speak like medical professionals, using organic healthcare conversations and simulated ones between patient actors and experienced nurses. This allows our system to express unique capabilities such as rapport building, trust building, empathy and bedside manner. Finally, we present the first comprehensive clinician evaluation of an LLM system for healthcare. We recruited over 1100 U.S. licensed nurses and over 130 U.S. licensed physicians to perform end-to-end conversational evaluations of our system by posing as patients and rating the system on several measures. We demonstrate Polaris performs on par with human nurses on aggregate across dimensions such as medical safety, clinical readiness, conversational quality, and bedside manner. Additionally, we conduct a challenging task-based evaluation of the individual specialist support agents, where we demonstrate our LLM agents significantly outperform a much larger general-purpose LLM (GPT-4) as well as from its own medium-size class (LLaMA-2 70B).

연구 동기 및 목표

  • 실시간으로 환자 대면 의료 대화 시스템 개발 안전성 강조 및 비진단적 지원 업무.
  • 다중 에이전트 구성으로 특화된 지원 모델을 통해 환각 감소 및 의료 정확도 향상.
  • AI 주도 대화에서 간호사 같은 친밀감, 공감, 간병 태도 가능하게.
  • 장기간 다-turn 상호작용에서 대화 상태 유지 및 음성 기반 커뮤니케이션 처리.
  • 시스템 성능을 인간 간호사와 비교 평가하고 전문 에이전트를 일반 LLM과 비교.

제안 방법

  • 상태를 가진 주요 에이전트와 여러 전문 지원 에이전트를 갖춘 다중 에이전트 LLM 구성 설계.
  • 주요 에이전트를 일반 지시 조정, 대화/에이전트 조정, 그리고 독점 의료 데이터 및 시뮬레이션 대화를 사용한 RLHF로 훈련.
  • 전문 에이전트들의 출력을 통해 대화 상태를 업데이트하는 메시지 전달 오케스트레이션 프레임워크 개발.
  • 프라이버시, 약물, 검사/생체 징후, 영양, 정책, EHR, 사람 개입 등 맥락과 안전 점검을 제공하는 전용 전문 에이전트 활용.
  • bf16/int8 정밀도, GQA, Flash Attention 2, RoPE 등 성능 최적화를 포함한 모델 및 데이터 전략 적용.
  • 시뮬레이션 환자-배우 대화 및 간호사 루프 데이터 생성을 활용하여 bedside manner 및 의학적 추론에 시스템을 훈련하고 정렬.

실험 결과

연구 질문

  • RQ1안전 중심 LLM 구성으로 실시간 대화 중 의료 안전, 임상 준비, 환자 교육, 대화 품질, bedside manner에서 인간 간호사와 동등해질 수 있는가?
  • RQ2전문 에이전트가 보건의료 특화 작업과 시뮬레이션에서 더 큰 일반 목적 LLM보다 우수한가?
  • RQ3작업별 모듈화가 실시간 간호 도우미 대화의 안전성, 지연, 정확도에 어떻게 영향을 미치는가?
  • RQ4간호사 같은 공감 및 의학 추론에 주된 대화 에이전트의 학습 및 데이터 전략은 무엇인가?
  • RQ5임상 대화에서 다중 에이전트 조정의 운영 상 도전과 안전 거래는 무엇인가?

주요 결과

  • Polaris 시스템은 의료 안전, 임상 준비, 환자 교육, 대화 품질, bedside manner의 집계 지표에서 인간 간호사와 동등한 성과를 보인다.
  • 전문 에이전트는 더 큰 일반 목적 LLM(GPT-4) 및 비슷한 크기(LLaMA-2 70B)보다 보건의료 작업에서 크게 우수.
  • 구성 아키텍처는 중복성, 전문화, 모듈식 업그레이드를 통해 안전하고 유지보수가 용이한 업데이트를 가능하게 하여 전체 시스템 재훈련 없이도 안전성 제공.
  • 필요 시 인간 감독이 있는 능동적 안전 가드레일을 호출하여 고위험 대화에서 안전성 강화.
  • 주요 에이전트는 대화의 유창성과 공감에 집중하고, 전문 에이전트는 정보를 검증하고 도메인 특화 작업을 관리하며 긴 대화에서 상태를 유지.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.