[論文レビュー] CXR-LLAVA: a multimodal large language model for interpreting chest X-ray images
CXR-LLAVA は、視覚変換器と大規模言語モデルを統合して胸部X線を解釈できるオープンソースのマルチモーダル大規模言語モデルです。592,580枚のCXRと自由記述型レントゲン報告書を用いて訓練・微調整し、内部テストではF1スコア0.81、外部テストではF1スコア0.62を達成。GPT-4-vision や Gemini-Pro-Vision を上回り、人間のレントゲン専門医による自律的報告で72.7%の成功率を示しました。
Purpose: This study aimed to develop an open-source multimodal large language model (CXR-LLAVA) for interpreting chest X-ray images (CXRs), leveraging recent advances in large language models (LLMs) to potentially replicate the image interpretation skills of human radiologists Materials and Methods: For training, we collected 592,580 publicly available CXRs, of which 374,881 had labels for certain radiographic abnormalities (Dataset 1) and 217,699 provided free-text radiology reports (Dataset 2). After pre-training a vision transformer with Dataset 1, we integrated it with an LLM influenced by the LLAVA network. Then, the model was fine-tuned, primarily using Dataset 2. The model's diagnostic performance for major pathological findings was evaluated, along with the acceptability of radiologic reports by human radiologists, to gauge its potential for autonomous reporting. Results: The model demonstrated impressive performance in test sets, achieving an average F1 score of 0.81 for six major pathological findings in the MIMIC internal test set and 0.62 for seven major pathological findings in the external test set. The model's F1 scores surpassed those of GPT-4-vision and Gemini-Pro-Vision in both test sets. In human radiologist evaluations of the external test set, the model achieved a 72.7% success rate in autonomous reporting, slightly below the 84.0% rate of ground truth reports. Conclusion: This study highlights the significant potential of multimodal LLMs for CXR interpretation, while also acknowledging the performance limitations. Despite these challenges, we believe that making our model open-source will catalyze further research, expanding its effectiveness and applicability in various clinical contexts. CXR-LLAVA is available at https://github.com/ECOFRI/CXR_LLAVA.
研究の動機と目的
- レントゲン専門医並みの性能で胸部X線画像を解釈できるマルチモーダル大規模言語モデルの開発。
- ラベル付き異常と自由記述型レントゲン報告書を併用した公的CXRデータセットを活用しての訓練。
- LLMを用いた自動レントゲン報告における診断精度とレポート品質の向上。
- 医療画像AI分野の研究と臨床応用を促進するオープンソースモデルの構築。
- 実世界のテスト環境において、最先端のビジョン・ランゲージモデルと比較して本モデルの性能を評価すること。
提案手法
- ラベル付き画像異常を有する374,881枚のCXRを用いて、ビジョン変換器を事前学習(データセット1)。
- LLAVAアーキテクチャを模倣した大規模言語モデルとビジョンエンコーダーを統合し、マルチモーダル基盤モデルを構築。
- 自由記述型レントゲン報告書を有する217,699枚のCXRを主なデータとして、融合モデルを微調整(データセット2)。
- 事前学習段階で視覚的・言語的表現の整合性を高めるために、対照的学習の目的関数を用いた。
- 微調整段階でビジョン・ランゲージのインstructチューニング戦略を採用し、レポート生成と診断推論の精度を向上。
- 人間のレントゲン専門医による評価を用いて、診断分類とレポート品質の両面でモデルを評価。
実験結果
リサーチクエスチョン
- RQ1多様なCXRデータで訓練されたマルチモーダルLLMは、人間のレントゲン専門医と同等の診断性能を達成できるか?
- RQ2CXR-LLAVAは、GPT-4-vision や Gemini-Pro-Vision といった最先端のビジョン・ランゲージモデルと比較して、CXR解釈タスクでどのように差をつけるか?
- RQ3CXR-LLAVAは、臨床的に受け入れられる自律的レポートをどの程度の割合で生成できるか?
- RQ4ラベル付き異常と自由記述型報告書の両方を訓練に用いることで、マルチモーダル医療LLMの性能にどのような影響を与えるか?
- RQ5CXR-LLAVAのオープンソース化は、医療画像AI分野におけるイノベーションを促進できるか?
主な発見
- CXR-LLAVAは、MIMICの内部テストセットにおける6つの主要病理所見で平均F1スコア0.81を達成した。
- 外部テストセットでは、7つの主要病理所見でF1スコア0.62を達成し、GPT-4-vision や Gemini-Pro-Vision を上回った。
- 人間のレントゲン専門医による評価において、自律的レポート生成の成功率は72.7%であった。これは、正解レポート(84.0%)よりは低いが、高い水準を示した。
- 内部および外部テストセットの両方で、CXR-LLAVAの性能はGPT-4-vision や Gemini-Pro-Vision を常に上回った。
- LLAVAを模倣したアーキテクチャを用いてビジョン変換器と大規模言語モデルを統合することで、CXR解釈において強力なゼロショットおよびフェイワショット一般化が実現された。
- CXR-LLAVAのオープンソース化により、医療画像AI分野におけるさらなる研究開発と臨床応用が促進されると予想される。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。