Skip to main content
QUICK REVIEW

[論文レビュー] BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs

Sheng Zhang, Yanbo Xu|arXiv (Cornell University)|Mar 2, 2023
Topic Modeling被引用数 95
ひとこと要約

BiomedCLIPはPMC-15M、PubMed Centralの15百万の画像キャプションデータセット上でドメイン特化型のビジョン-ランゲージ基盤モデルを事前学習させ、生物医学領域における検索、分類、VQAで最先端の成果を達成する。

ABSTRACT

Biomedical data is inherently multimodal, comprising physical measurements and natural language narratives. A generalist biomedical AI model needs to simultaneously process different modalities of data, including text and images. Therefore, training an effective generalist biomedical model requires high-quality multimodal data, such as parallel image-text pairs. Here, we present PMC-15M, a novel dataset that is two orders of magnitude larger than existing biomedical multimodal datasets such as MIMIC-CXR, and spans a diverse range of biomedical image types. PMC-15M contains 15 million biomedical image-text pairs collected from 4.4 million scientific articles. Based on PMC-15M, we have pretrained BiomedCLIP, a multimodal foundation model, with domain-specific adaptations tailored to biomedical vision-language processing. We conducted extensive experiments and ablation studies on standard biomedical imaging tasks from retrieval to classification to visual question-answering (VQA). BiomedCLIP achieved new state-of-the-art results in a wide range of standard datasets, substantially outperforming prior approaches. Intriguingly, by large-scale pretraining on diverse biomedical image types, BiomedCLIP even outperforms state-of-the-art radiology-specific models such as BioViL in radiology-specific tasks such as RSNA pneumonia detection. In summary, BiomedCLIP is a fully open-access foundation model that achieves state-of-the-art performance on various biomedical tasks, paving the way for transformative multimodal biomedical discovery and applications. We release our models at https://aka.ms/biomedclip to facilitate future research in multimodal biomedical AI.

研究の動機と目的

  • 大規模で多様性があり、オープンな生物医学ビジョン-ランゲージデータセットの必要性に対処する。
  • 長い生物医学的キャプションと高解像度画像を活用したドメイン特化型マルチモーダルモデルを事前学習する。
  • BiomedCLIPを検索、ゼロショット分類、医療VQAの各領域で評価し、最先端の性能を確立する。
  • 大規模で多様な生物医学_pretrainingが放射線科中心のモデルを特定のタスクで上回ることを示す。

提案手法

  • PubMed Centralの記事から15百万の図とキャプションの公開データセットであるPMC-15Mを作成する。
  • テキストエンコーダを生物医学向けに調整したPubMedBERTとより大きなビジョンエンコーダを高解像度で用い、CLIPを適応させる。
  • パッチ dropoutと適切なバッチサイズを実装し、生物医学事前学習の効率と性能を最適化する。
  • PMC-Fine-Grained-46Mを構築してデータの多様性を高め、分割パネルやインラインcitancesを含むより細粒度な画像-テキスト対を作成する。
  • 検索、分類、VQAを含む8つの標準的な生物医学ビジョン-ランゲージタスクを用いて評価し、一般ドメインのCLIP、PubMedCLIP、MedCLIP、BioViLと比較する。

実験結果

リサーチクエスチョン

  • RQ1大規模で多様なオープンな生物医学画像-テキストデータセットで訓練すると、生物医学分野のビジョン-ランゲージの汎用性は向上するのか?
  • RQ2BiomedCLIPは検索、分類、VQAタスクで一般ドメインモデルおよび放射線科に焦点を当てたモデルとどう比較されるのか?
  • RQ3生物医学のマルチモーダル性能を最大化するためのドメイン特有の適応(テキストおよび画像エンコーダ、トークン化、画像解像度)は何か?
  • RQ4多様な事前学習データは生物医学の分野間での転移(例:放射線科 vs 病理学)を可能にするのか?

主な発見

  • BiomedCLIPはクロスモーダル検索精度が高く、Top-1およびTop-5リコールが一般ドメインのCLIPおよびPubMedCLIPを大幅に上回る。
  • BiomedCLIPは5つのデータセットにわたるゼロショット画像分類で優れた性能を示し、RSNAベンチマークで放射線科向けのBioViLよりもラベル付きデータが少なくても上回る。
  • BiomedCLIPは医療VQAベンチマーク(VQ-RAD、SLAKE)で最先端の結果を達成し、オープンエンドの質問にも堅牢な性能を示す。
  • PMC-15Mの多様な画像タイプでの事前学習は、放射線科のみの事前学習をいくつかのタスクで超える堅牢な表現をもたらす。
  • この研究は、大規模でドメイン特化型のマルチモーダル事前学習がオープンアクセスで高性能な生物医学AIモデルを実現できることを示している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。