Skip to main content
QUICK REVIEW

[論文レビュー] MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models

Yan Cai, Linlin Wang|arXiv (Cornell University)|Dec 20, 2023
Topic Modeling被引用数 4
ひとこと要約

MedBenchは、公式試験および臨床症例から得た40,041件の中国語医療質問から構成される大規模かつマルチソースのベンチマークであり、中国語医療LLMの知識および推論能力を中国語医療教育の全範囲にわたって評価することを目的としている。この研究では、中国語医療LLMが顕著に劣っていることが明らかになった一方で、ChatGPTのような一般ドメインモデルは強力な医療知識を示しており、臨床的推論および診断の正確性における格差が浮き彫りになった。

ABSTRACT

The emergence of various medical large language models (LLMs) in the medical domain has highlighted the need for unified evaluation standards, as manual evaluation of LLMs proves to be time-consuming and labor-intensive. To address this issue, we introduce MedBench, a comprehensive benchmark for the Chinese medical domain, comprising 40,041 questions sourced from authentic examination exercises and medical reports of diverse branches of medicine. In particular, this benchmark is composed of four key components: the Chinese Medical Licensing Examination, the Resident Standardization Training Examination, the Doctor In-Charge Qualification Examination, and real-world clinic cases encompassing examinations, diagnoses, and treatments. MedBench replicates the educational progression and clinical practice experiences of doctors in Mainland China, thereby establishing itself as a credible benchmark for assessing the mastery of knowledge and reasoning abilities in medical language learning models. We perform extensive experiments and conduct an in-depth analysis from diverse perspectives, which culminate in the following findings: (1) Chinese medical LLMs underperform on this benchmark, highlighting the need for significant advances in clinical knowledge and diagnostic precision. (2) Several general-domain LLMs surprisingly possess considerable medical knowledge. These findings elucidate both the capabilities and limitations of LLMs within the context of MedBench, with the ultimate goal of aiding the medical research community.

研究の動機と目的

  • 中国語医療LLMの標準的で大規模な評価ベンチマークが、現実の臨床および教育的基準を反映していないという問題を是正すること。
  • 中国医師の教育および臨床の全プロセスを反映する、中国医師資格試験、研修医標準化試験、主任医師資格試験といった多様で本物の出題源を含めることで、既存ベンチマークの限界を克服すること。
  • 電子カルテから得た本物の臨床症例を統合し、教科書的知識を超えた実践的診断的推論能力を評価すること。
  • 項目反応理論(IRT)を用いた心理測定的妥当性のあるベンチマークを提供することで、信頼性の高いスケーラブルなモデル評価を実現すること。
  • 中国語特化型および一般ドメインLLMの臨床的知識想起、推論、説明生成における強みと弱みを特定すること。

提案手法

  • 中国医師資格試験、研修医標準化試験、主任医師資格試験、および本物の臨床症例から得た40,041件の質問を統合し、MedBenchを構築した。
  • 内部科、外科学、婦人科、小児科の4つの医学分野に分類し、レベル1からレベル9までの階層的難易度スケールを用いて構成した。
  • 単一ホップ、文の特定、マルチホップ推論の複数の推論タイプを統合し、論理的および診断的推論能力を評価した。
  • 回答の根拠となる説明をモデル出力として収集し、推論プロセスの整合性および正しさを評価した。
  • 項目反応理論(IRT)を用いて質問の難易度およびモデルのパフォーマンスをモデル化し、より信頼性が高くスケーラブルな評価を可能にした。
  • 人間による評価を実施し、モデル出力の妥当性、会話の質および臨床的推論の正確性を評価した。
Figure 1: Comparison of procedures in different countries.
Figure 1: Comparison of procedures in different countries.

実験結果

リサーチクエスチョン

  • RQ1中国語医療LLMは、中華人民共和国の医師が経験する教育的および臨床的全プロセスを反映する包括的ベンチマークにおいて、どのように性能を発揮するか?
  • RQ2ChatGPTのような一般ドメインLLMは、ドメイン特化型モデルと比較して、中国語文脈においてどの程度医療知識を移転可能としているか?
  • RQ3LLMは回答の論理的根拠をどの程度適切に生成できるか?説明の質は回答の正答率と相関しているか?
  • RQ4現在の中国語医療LLMにおいて、診断的推論および臨床的知識想起における主な制限要因は何か?
  • RQ5項目反応理論(IRT)のような心理測定的手法は、医療LLMの評価の信頼性およびスケーラビリティを向上させることができるか?

主な発見

  • 中国語医療LLMはMedBenchにおいて顕著に劣っており、臨床的知識および診断の正確性の向上に大きな余地があることが示された。
  • ChatGPTを含む複数の一般ドメインLLMは、強力な医療知識および推論能力を示しており、一般ドメインから医療ドメインへの知識移転の可能性が示唆された。
  • 人間による評価では、ChatGPTは豊富な臨床的知識を示したが、現在の中国語医療LLMは高品質な会話能力および十分な事実の正確性に欠けていることが判明した。
  • 誤った回答を出す際、LLMはしばしば非論理的な根拠を提示しており、知識の欠落または誤った推論メカニズムの両方が原因である可能性がある。
  • ChatGPTが複雑な症例分析の質問で劣悪なパフォーマンスを示したことは、深い医療理解および臨床的推論能力の欠如という深刻な欠陥を示している。
  • 項目反応理論(IRT)の統合により、より信頼性が高くスケーラブルなモデル評価が可能となり、今後の評価手法の洗練に貢献できることが明らかになった。
Figure 2: Examples of prompts and corresponding answers. The left side shows the prompt and responses of LLMs on an example question from an exam. The right side is an example of a real-world case.
Figure 2: Examples of prompts and corresponding answers. The left side shows the prompt and responses of LLMs on an example question from an exam. The right side is an example of a real-world case.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。