Skip to main content
QUICK REVIEW

[論文レビュー] CMB: A Comprehensive Medical Benchmark in Chinese

Xidong Wang, Guiming Hardy Chen|arXiv (Cornell University)|Aug 17, 2023
Machine Learning in Healthcare被引用数 4
ひとこと要約

本論文は、中国独自の言語的・文化的医療文脈を反映した、大規模言語モデル(LLMs)の評価を目的とした包括的で中国語向けの医療ベンチマーク、CMBを紹介する。CMBは、本物の中国語医師国家資格試験問題と臨床事例に基づいて構築されており、知識の想起(CMB-Exam)と臨床的推論(CMB-Clin)の両方を評価する。自動評価(GPT-4)と専門家評価の間で強い整合性を示しており、スピアマン相関係数は0.93に達する。

ABSTRACT

Large Language Models (LLMs) provide a possibility to make a great breakthrough in medicine. The establishment of a standardized medical benchmark becomes a fundamental cornerstone to measure progression. However, medical environments in different regions have their local characteristics, e.g., the ubiquity and significance of traditional Chinese medicine within China. Therefore, merely translating English-based medical evaluation may result in extit{contextual incongruities} to a local region. To solve the issue, we propose a localized medical benchmark called CMB, a Comprehensive Medical Benchmark in Chinese, designed and rooted entirely within the native Chinese linguistic and cultural framework. While traditional Chinese medicine is integral to this evaluation, it does not constitute its entirety. Using this benchmark, we have evaluated several prominent large-scale LLMs, including ChatGPT, GPT-4, dedicated Chinese LLMs, and LLMs specialized in the medical domain. We hope this benchmark provide first-hand experience in existing LLMs for medicine and also facilitate the widespread adoption and enhancement of medical LLMs within China. Our data and code are publicly available at https://github.com/FreedomIntelligence/CMB.

研究の動機と目的

  • 英語中心の医療ベンチマークが中国独自の医療文脈におけるLLMsの評価に限界を示す問題に対処すること。
  • 中国の医学教育と臨床実務、特に伝統的中国医学を反映した、文化的に根ざした標準化された評価フレームワークを構築すること。
  • 競争的リーダーボードではなく、自己評価ツールとしての役割を果たし、医療応用分野における継続的モデル改善を可能にすること。
  • 本物の医師国家資格試験データを用いて評価の厳密性を確保し、客観的な真値を提供するとともに、専門家水準のパフォーマンスを示す基準として60%を設定すること。
  • GPT-4を用いた自動評価の信頼性を、臨床的推論タスクにおける専門家人間の判断と比較することで検証すること。

提案手法

  • CMBは、医師、看護師、医療技術者、薬剤師の全分野をカバーする、本物の中国語医師国家資格試験問題から構築される。
  • CMB-Examは、4〜6つの選択肢と1つ以上の正解を有する多肢選択問題から構成され、大学部から上級専門職称号までの公式試験問題から抽出される。
  • CMB-Clinは、医学的知識と推論能力の統合を要する複雑な、複数ターンの臨床ケーススタディから成り、回答は専門家によるコンSENSUSで検証される。
  • 公に利用可能な匿名化された試験データを用い、中国医療問題データベースの明示的承認を得ることで、倫理的適合性とデータプライバシーを確保する。
  • 自動評価は、GPT-4を用いて、関連性、包括性、流暢さ、熟練度の観点からモデル出力をスコア付けし、専門家による人間のアノテーションと比較する。
  • 統計的妥当性検証には、GPT-4評価と専門家ランク付けの間のスピアマン相関(0.93)と、デコードハイパーパrameter(例:温度)がパフォーマンス安定性に与える影響の分析を含む。
Figure 1 : Components of the CMB dataset. Left: The structure of CMB-Exam, consisting of multiple-choice and multiple-answer questions. Right: an example of CMB-Clin. Each example consists of a description and a multi-turn conversation.
Figure 1 : Components of the CMB dataset. Left: The structure of CMB-Exam, consisting of multiple-choice and multiple-answer questions. Right: an example of CMB-Clin. Each example consists of a description and a multi-turn conversation.

実験結果

リサーチクエスチョン

  • RQ1最先端のLLMsは、中国語で構築された包括的で文化的・言語的にローカライズされた医療ベンチマークで、どの程度のパフォーマンスを示すか?
  • RQ2GPT-4を用いた自動評価は、中国語医療分野の臨床的推論タスクにおいて、専門家の人間評価とどの程度整合するか?
  • RQ3知識集約型の多肢選択問題(CMB-Exam)のパフォーマンスは、複雑な臨床的推論タスク(CMB-Clin)のパフォーマンスを予測できるか?
  • RQ4推論時のデコード温度の変更に対して、モデルランク付けはどの程度頑健か?
  • RQ5CMB-Examデータでのファインチューニングは、モデルの臨床ケースベース推論能力への一般化能力にどのような影響を与えるか?

主な発見

  • GPT-4、ChatGPT、Baichuan-13B-chatは、CMB-Clinでそれぞれ平均3.45、3.40、3.74のスコアを記録し、関連性・包括性・熟練度の観点で他のモデルと比較して最低7.4%以上優れていた。
  • CMB-ClinにおけるGPT-4評価と専門家人間評価の間のスピアマン相関は0.93であり、自動評価の高い整合性と信頼性を示している。
  • CMB-ExamとCMB-Clinにおけるモデルランク付けの相関は0.89(p値 = 2.3e-4)であり、知識想起と推論タスクの両方で一貫したパフォーマンスを示していることが示された。
  • CMB-Clinにおけるパフォーマンスは、デコード温度の上昇に伴い低下し、温度0で最高スコアを記録した。これは、医療文脈では決定論的出力が好まれることを示唆している。
  • 異なる温度設定におけるモデルランク付けのペairワイズスピアマン相関は、0.87以上を維持しており、モデルランク付けがハイパーパrameterの変更に対して頑健であることが示された。
  • MedicalGPT、DoctorGLM、Bentsao、ChatGLM-Medなどのモデルは、流暢さを除いて満足のいく結果を示さず、医療知識や推論能力の欠落が原因である可能性がある。
Figure 2 : Accuracy across various clinical medicine fields at different career stages. The accuracies are the Zero-shot average values for TOP-5 models using direct response strategy.
Figure 2 : Accuracy across various clinical medicine fields at different career stages. The accuracies are the Zero-shot average values for TOP-5 models using direct response strategy.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。