Skip to main content
QUICK REVIEW

[論文レビュー] Polaris: A Safety-focused LLM Constellation Architecture for Healthcare

Subhabrata Mukherjee, Paul Gamble|arXiv (Cornell University)|Mar 20, 2024
Quality and Safety in Healthcare被引用数 19
ひとこと要約

Polaris は医療安全を重視したリアルタイムの患者向け対話のためのマルチエージェント LLM コンステレーションを提供し、主エージェントと専門サポートエージェントを用いて医療安全性を高め、幻覚を減らします。

ABSTRACT

We develop Polaris, the first safety-focused LLM constellation for real-time patient-AI healthcare conversations. Unlike prior LLM works in healthcare focusing on tasks like question answering, our work specifically focuses on long multi-turn voice conversations. Our one-trillion parameter constellation system is composed of several multibillion parameter LLMs as co-operative agents: a stateful primary agent that focuses on driving an engaging conversation and several specialist support agents focused on healthcare tasks performed by nurses to increase safety and reduce hallucinations. We develop a sophisticated training protocol for iterative co-training of the agents that optimize for diverse objectives. We train our models on proprietary data, clinical care plans, healthcare regulatory documents, medical manuals, and other medical reasoning documents. We align our models to speak like medical professionals, using organic healthcare conversations and simulated ones between patient actors and experienced nurses. This allows our system to express unique capabilities such as rapport building, trust building, empathy and bedside manner. Finally, we present the first comprehensive clinician evaluation of an LLM system for healthcare. We recruited over 1100 U.S. licensed nurses and over 130 U.S. licensed physicians to perform end-to-end conversational evaluations of our system by posing as patients and rating the system on several measures. We demonstrate Polaris performs on par with human nurses on aggregate across dimensions such as medical safety, clinical readiness, conversational quality, and bedside manner. Additionally, we conduct a challenging task-based evaluation of the individual specialist support agents, where we demonstrate our LLM agents significantly outperform a much larger general-purpose LLM (GPT-4) as well as from its own medium-size class (LLaMA-2 70B).

研究の動機と目的

  • リアルタイムの患者向け医療対話システムを開発し、安全性と非診断的サポート作業を強調する。
  • 幻覚を減らし、専門サポートモデルを備えたマルチエージェント星座で医療の正確性を高める。
  • AI駆動対話で看護師のようなラポール、共感、ベッドサイドマナーを実現。
  • 長い多ターン対話で対話状態を維持し、音声ベースの通信を扱う。
  • 人間の看護師と比較してシステム性能を評価し、専門エージェントと一般的な LLM を比較する。

提案手法

  • 状態を持つ主エージェントと複数の専門サポートエージェントを備えたマルチエージェント LLM コンステレーションを設計する。
  • 独自の医療データと模擬対話を用いた一般指示調整、対話/エージェント調整、RLHFで主エージェントを訓練する。
  • 専門エージェントの出力を介して対話状態を更新するメッセージ伝達オーケストレーションフレームワークを開発する。
  • プライバシー、薬剤、検査/バイタル、栄養、ポリシー、EHR、人間介入 等の専用専門エージェントを使用して文脈と安全性チェックを提供する。
  • 50-100B パラメータのサポートモデル、bf16/int8 精度、パフォーマンス最適化(GQA、Flash Attention 2、RoPE)を含むモデルとデータ戦略を適用する。
  • シミュレートされた患者役の対話と看護師を組み込んだデータ生成を用いて、ベッドサイドのマナーと医療推論に合わせるためにシステムを訓練・整合させる。

実験結果

リサーチクエスチョン

  • RQ1安全性に焦点を当てた LLM コンステレーションは、リアルタイムの対話中に医療安全性、臨床準備、患者教育、対話品質、ベッドサイドマナーで人間の看護師と同等に達成できるか?
  • RQ2専門サポートエージェントは、医療特化のタスクとシミュレーションでより大きな一般目的 LLM(GPT-4)および同程度のサイズの LLM(LLaMA-2 70B)を上回るか?
  • RQ3タスク固有のモジュール化は、安全性、待ち時間、そしてリアルタイムの看護師支援対話における正確性にどう影響するか?
  • RQ4主対話エージェントと看護師のような共感と医療推論を最もよく整合させる訓練・データ戦略は何か?
  • RQ5臨床対話におけるマルチエージェント協調の運用上の課題と安全性のトレードオフは何か?

主な発見

  • Polaris システムは、エンドツーエンドの評価において、医療安全性、臨床準備、患者教育、対話品質、ベッドサイドマナーの総合指標で人間の看護師と同等である。
  • 専門サポートエージェントは、構成評価において、より大きな一般目的 LLM(GPT-4)および同程度のサイズの LLM(LLaMA-2 70B)を医療タスクで大幅に上回る。
  • 星座アーキテクチャは、冗長性、専門化、モジュラーアップグレードを通じて安全性の利点を提供し、システム全体を再訓練せずにより安全で保守可能な更新を可能にする。
  • 必要に応じて人間の監視下で能動的な安全ガードレールを起動でき、高リスク対話の安全性を高める。
  • 主エージェントは会話の流暢さと共感に焦点を当て、専門エージェントは情報検証、ドメイン特有のタスク管理、長い対話で状態を維持する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。