[論文レビュー] Evaluating LLM -- Generated Multimodal Diagnosis from Medical Images and Symptom Analysis
この論文は、マルチモーダルな相互作用とドメイン特有の分析を組み合わせた2段階のLLM評価パラダイムを提案し、画像付きの病理学MCQに対するGPT-4-Vision-Previewを評価し、約84%の精度を達成し、NERと知識グラフを用いて洞察を抽出する。
Large language models (LLMs) constitute a breakthrough state-of-the-art Artificial Intelligence technology which is rapidly evolving and promises to aid in medical diagnosis. However, the correctness and the accuracy of their returns has not yet been properly evaluated. In this work, we propose an LLM evaluation paradigm that incorporates two independent steps of a novel methodology, namely (1) multimodal LLM evaluation via structured interactions and (2) follow-up, domain-specific analysis based on data extracted via the previous interactions. Using this paradigm, (1) we evaluate the correctness and accuracy of LLM-generated medical diagnosis with publicly available multimodal multiple-choice questions(MCQs) in the domain of Pathology and (2) proceed to a systemic and comprehensive analysis of extracted results. We used GPT-4-Vision-Preview as the LLM to respond to complex, medical questions consisting of both images and text, and we explored a wide range of diseases, conditions, chemical compounds, and related entity types that are included in the vast knowledge domain of Pathology. GPT-4-Vision-Preview performed quite well, scoring approximately 84\% of correct diagnoses. Next, we further analyzed the findings of our work, following an analytical approach which included Image Metadata Analysis, Named Entity Recognition and Knowledge Graphs. Weaknesses of GPT-4-Vision-Preview were revealed on specific knowledge paths, leading to a further understanding of its shortcomings in specific areas. Our methodology and findings are not limited to the use of GPT-4-Vision-Preview, but a similar approach can be followed to evaluate the usefulness and accuracy of other LLMs and, thus, improve their use with further optimization.
研究の動機と目的
- 医療分野のマルチモーダルLLMを評価するための2段階の評価パラダイムを提案する(マルチモーダル評価とドメイン特有の分析)。
- 画像とテキストのMCQを用いてLLM生成の病理診断の正確性を評価する。
- 正答と誤答および説明からIMA、NER、知識グラフを用いて弱点を分析し微調整の指針を得る。
- GPT-4-Vision-Previewを用いて公開病理MCQで方法論を実証する。
提案手法
- 事前に定められた行動規範を用いた画像+テキストMCQによる構造化されたマルチモーダル対話。
- 特定の回答形式と簡潔な回答を強制するプロンプト設計。
- Image Metadata Analysis (IMA)、Named Entity Recognition (NER)、Knowledge Graphs (KGs)を含むデータ抽出。
- 正答/誤答と説明から微調整の要件を導くためのドメイン特化分析。

実験結果
リサーチクエスチョン
- RQ1医療画像と症状ベースのテキストを組み合わせて診断する際のLLMの正確さはどの程度か?
- RQ2IMA、NER、KG分析を通じてLLMが示す弱点や知識経路は何か?
- RQ3評価手法はドメイン特有の医療タスクのターゲットを絞った微調整や再訓練を指針づけられるか?
- RQ4このアプローチはGPT-4-Vision-Preview以外の他のLLMにも一般化可能か?
主な発見
- GPT-4-Vision-Previewはマルチモーダル病理MCQで約84%の正答を達成した。
- 正答: 66/79; 誤答: 13/79、79問の画像・質問アイテム全体で。
- 病理のサブドメインはパフォーマンスにばらつきがあり、例:General Pathology: Atherosclerosis & Thrombosis 8/10、Cell Injury 7/10、ImmunoPathology 9/10、Inflammation 6/10、Neoplasia 10/10;Organ System Pathology: Cardiovascular 9/9、DermatoPathology 9/10、Endocrine 8/10。
- IMAは画像領域の弱点を主に Cardiovascular、Skin、Endocrine 画像で特定した。
- NERとKG分析は微調整の標的となる特定のエンティティと知識経路の弱点を明らかにした。

より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。