Skip to main content
QUICK REVIEW

[論文レビュー] Evaluating multiple large language models in pediatric ophthalmology

Jason Holmes, Rui Peng|arXiv (Cornell University)|Nov 7, 2023
Artificial Intelligence in Healthcare and Education被引用数 4
ひとこと要約

本研究では、小児眼科分野における100問の多肢選択式試験を用いて、GPT-4、ChatGPT(GPT-3.5)、PaLM2という3つの大規模言語モデルの性能を、医学部学生、大学院生、および専門医の臨床医と比較した。GPT-4は専門医と同等の正確性を示し、医学部学生を上回り、他のモデルと比較して優れた応答の安定性と自信を示した。

ABSTRACT

IMPORTANCE The response effectiveness of different large language models (LLMs) and various individuals, including medical students, graduate students, and practicing physicians, in pediatric ophthalmology consultations, has not been clearly established yet. OBJECTIVE Design a 100-question exam based on pediatric ophthalmology to evaluate the performance of LLMs in highly specialized scenarios and compare them with the performance of medical students and physicians at different levels. DESIGN, SETTING, AND PARTICIPANTS This survey study assessed three LLMs, namely ChatGPT (GPT-3.5), GPT-4, and PaLM2, were assessed alongside three human cohorts: medical students, postgraduate students, and attending physicians, in their ability to answer questions related to pediatric ophthalmology. It was conducted by administering questionnaires in the form of test papers through the LLM network interface, with the valuable participation of volunteers. MAIN OUTCOMES AND MEASURES Mean scores of LLM and humans on 100 multiple-choice questions, as well as the answer stability, correlation, and response confidence of each LLM. RESULTS GPT-4 performed comparably to attending physicians, while ChatGPT (GPT-3.5) and PaLM2 outperformed medical students but slightly trailed behind postgraduate students. Furthermore, GPT-4 exhibited greater stability and confidence when responding to inquiries compared to ChatGPT (GPT-3.5) and PaLM2. CONCLUSIONS AND RELEVANCE Our results underscore the potential for LLMs to provide medical assistance in pediatric ophthalmology and suggest significant capacity to guide the education of medical students.

研究の動機と目的

  • 大規模言語モデル(LLM)の小児眼科分野における臨床的推論の正確性を評価すること。
  • LLMの性能を医学部学生、大学院生、および専門医と比較すること。
  • LLMと人間の被験者間での応答の安定性、自信、相関関係を評価すること。
  • 専門的医療分野におけるLLMの教育的および臨床的支援の可能性を検討すること。

提案手法

  • 小児眼科分野の内容に基づいて、100問の多肢選択式試験が作成された。
  • ChatGPT(GPT-3.5)、GPT-4、PaLM2という3つのLLMが、ネットワークインターフェース経由で試験問題に回答した。
  • 医学部学生、大学院生、および専門医という3つの人間の被験者グループが、同じ試験を完了した。
  • 成績は平均得点、応答の安定性、自信度、人間の回答との相関関係によって測定された。
  • 統計解析により、LLMと人間のグループが主な指標で比較された。

実験結果

リサーチクエスチョン

  • RQ1異なるLLMは、小児眼科分野における人間の臨床医と比較してどのように性能を発揮するか?
  • RQ2LLMは医学部学生や専門医の診断正確性に匹敵または上回ることができるか?
  • RQ3LLMの応答は、人間の応答と比較して、どれほど安定的で、自信があるとされるか?
  • RQ4LLMと人間の成績の間には、臨床的推論タスクにおいてどの程度の相関関係があるか?

主な発見

  • GPT-4は専門医と同等の平均得点を達成し、小児眼科分野における高い臨床的推論の正確性を示した。
  • ChatGPT(GPT-3.5)およびPaLM2は医学部学生を上回ったが、大学院生の成績をわずかに下回った。
  • GPT-4は、ChatGPT(GPT-3.5)およびPaLM2と比較して、より高い応答の安定性と自信度スコアを示した。
  • GPT-4と専門医の成績の相関関係は、他のLLMと人間のペアよりも強かった。
  • LLM、特にGPT-4は、100問のテストセット全体で一貫性があり、信頼性の高い性能を示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。