Skip to main content
QUICK REVIEW

[論文レビュー] BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks

Kai Zhang, Rong Zhou|arXiv (Cornell University)|May 26, 2023
Artificial Intelligence in Healthcare and Education参考文献 84被引用数 51
ひとこと要約

BiomedGPT は、多様な生物医療タスクに対するオープンソースの汎用ビジョン-言語モデルで、26データセット、25実験中15件でSOTAを達成し、182Mパラメータでゼロショット転送を可能にします。

ABSTRACT

Traditional biomedical artificial intelligence (AI) models, designed for specific tasks or modalities, often exhibit limited flexibility in real-world deployment and struggle to utilize holistic information. Generalist AI holds the potential to address these limitations due to its versatility in interpreting different data types and generating tailored outputs for diverse needs. However, existing biomedical generalist AI solutions are typically heavyweight and closed source to researchers, practitioners, and patients. Here, we propose BiomedGPT, the first open-source and lightweight vision-language foundation model, designed as a generalist capable of performing various biomedical tasks. BiomedGPT achieved state-of-the-art results in 16 out of 25 experiments while maintaining a computing-friendly model scale. We also conducted human evaluations to assess the capabilities of BiomedGPT in radiology visual question answering, report generation, and summarization. BiomedGPT exhibits robust prediction ability with a low error rate of 3.8% in question answering, satisfactory performance with an error rate of 8.3% in writing complex radiology reports, and competitive summarization ability with a nearly equivalent preference score to human experts. Our method demonstrates that effective training with diverse data can lead to more practical biomedical AI for improving diagnosis and workflow efficiency.

研究の動機と目的

  • 画像とテキストにまたがる多様な生物医療タスクに対して、統一された汎用AIモデルを構想する。
  • 多様な多モーダル生物医療データで単一モデルを事前学習させ、普遍的な表現を学習する。
  • 視覚のみ、言語、およびマルチモーダルタスクを処理するために、タスク固有の指示でファインチューニングする。
  • 大規模なクローズドモデルと比較して性能を評価し、実世界での展開可能性を示す。
  • モデルチェックポイント、トレーニングデータの処理、コードをオープンソース化することによって透明性を促進する。

提案手法

  • BERT風のエンコーダとGPT風のデコーダを備えた、シーケンスツーシーケンスの BiomedGPT を用いる。
  • 画像、テキスト、マルチモーダルデータを網羅する14の自由利用可能な生物医療データセットで事前学習する(352,567枚の画像; 約183M テキスト文; 46,408のオブジェクトラベルペア; 約271,803の画像-テキストペア)。
  • 5つの事前学習タスクを用いる:画像のみのマスク付き画像モデリング、テキストのみのマスク付き言語モデリング、画像キャプション生成、ビジュアル質問応答(VQA)を含む。
  • 指示ベースのプロンプトでファインチューニングを行い、5つの医療AIタスク(分類、言語理解、要約、キャプション生成、VQA)にまたがる25の下流データセットに適用。
  • モデル規模3種(BiomedGPT-S、-M、-B)を探索し、スケール効果を研究する;一般知識と医療知識を注入するためにOFAから初期化する。
Figure 1: The overview of BiomedGPT: workflow, performance and pretraining datasets. (a) Graphical illustration of how BiomedGPT handle multi-modal inputs and perform diverse downstream tasks. The expected form of output for each task is determined by feeding the specific instruction to the model. (
Figure 1: The overview of BiomedGPT: workflow, performance and pretraining datasets. (a) Graphical illustration of how BiomedGPT handle multi-modal inputs and perform diverse downstream tasks. The expected form of output for each task is determined by feeding the specific instruction to the model. (

実験結果

リサーチクエスチョン

  • RQ1視覚と言語の両方をカバーする複数の生物医療モダリティとタスクを、単一の統一モデルで効果的に処理できるか。
  • RQ2多様でマルチタスクな事前学習は、生物医学における一般化や下流パフォーマンスを改善するか。
  • RQ3モデルのサイズは、視覚、言語理解、およびマルチモーダルな生物医療タスクのパフォーマンスにどう影響するか。
  • RQ4統一された生物医療モデルのゼロショットVQA能力は、大規模マルチモーダルシステムと比べてどうか。
  • RQ5BiomedGPT は実世界の展開と放射線科医の評価でどのように性能を発揮するか。

主な発見

  • BiomedGPT は、26データセットにわたる25実験のうち15件でSOTAを達成。
  • 182Mパラメータの BiomedGPT-B モデルは、VQA-RAD および SLAKE で Med-PaLM M (12B) を大幅に上回り、PathVQA にほぼ匹敵。
  • 画像キャプション生成では、BiomedGPT が PEIR GROSS でSOTAを上回り、CIDErの大幅な向上を示す;CIDErの改善はデータセット全体で強調される。
  • 11個の MedMNIST由来データセットにわたる医用画像分類で、モデルサイズの拡大とともに高解像度データセットで顕著な改善を見せ、いくつかのベースラインを上回る。
  • BiomedGPT は強力なゼロショットVQA能力を示し、追加のタスク固有要素なしで自由形式の回答を生成し、GPT-4V の一部評価で競合的な性能を達成する。
Figure 2: BiomedGPT performs fine-tuning for vision-language downstream tasks. (a) Graphical illustration of inference workflow of BiomedGPT for VQA task. Our model can discrete both visual and linguistic inputs from questions into tokens and generated the corresponding answers. (b) VQA performance
Figure 2: BiomedGPT performs fine-tuning for vision-language downstream tasks. (a) Graphical illustration of inference workflow of BiomedGPT for VQA task. Our model can discrete both visual and linguistic inputs from questions into tokens and generated the corresponding answers. (b) VQA performance

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。