Skip to main content
QUICK REVIEW

[論文レビュー] FairMedFM: Fairness Benchmarking for Medical Imaging Foundation Models

Ruinan Jin, Zikang Xu|arXiv (Cornell University)|Jul 1, 2024
Health Systems, Economic Evaluations, Quality of Life被引用数 4
ひとこと要約

FairMedFMは、画像診断分野の基盤モデル(FMs)における公平性を包括的に評価するベンチマークを提示する。17の多様な医療画像データセット(モダリティや次元が異なる)を統合し、分類およびセグメンテーションの複数の下流タスクと公平性に配慮した手法を用いて20のFMsを評価している。主な発見として、モデル間で一貫したバイアスが継続的に存在し、有用性と公平性のトレードオフが顕著で、既存のバイアス低減戦略の有効性は限定的であり、どのFMsも公平性または有用性において一貫して他のモデルを上回ることはなかった。

ABSTRACT

The advent of foundation models (FMs) in healthcare offers unprecedented opportunities to enhance medical diagnostics through automated classification and segmentation tasks. However, these models also raise significant concerns about their fairness, especially when applied to diverse and underrepresented populations in healthcare applications. Currently, there is a lack of comprehensive benchmarks, standardized pipelines, and easily adaptable libraries to evaluate and understand the fairness performance of FMs in medical imaging, leading to considerable challenges in formulating and implementing solutions that ensure equitable outcomes across diverse patient populations. To fill this gap, we introduce FairMedFM, a fairness benchmark for FM research in medical imaging.FairMedFM integrates with 17 popular medical imaging datasets, encompassing different modalities, dimensionalities, and sensitive attributes. It explores 20 widely used FMs, with various usages such as zero-shot learning, linear probing, parameter-efficient fine-tuning, and prompting in various downstream tasks -- classification and segmentation. Our exhaustive analysis evaluates the fairness performance over different evaluation metrics from multiple perspectives, revealing the existence of bias, varied utility-fairness trade-offs on different FMs, consistent disparities on the same datasets regardless FMs, and limited effectiveness of existing unfairness mitigation methods. Checkout FairMedFM's project page and open-sourced codebase, which supports extendible functionalities and applications as well as inclusive for studies on FMs in medical imaging over the long term.

研究の動機と目的

  • 医療画像分野の基盤モデル(FMs)における公平性を評価するための標準的で包括的なベンチマークが不足しているという問題に対処すること。
  • 多様なFMs、データモダリティ、感受性属性(例:性別、年齢、人種)および下流タスク(分類およびセグメンテーション)を網羅的に評価することで、公平性を体系的に分析すること。
  • さまざまなファインチューニング戦略およびFMsの種別に応じた、モデルの有用性と公平性のトレードオフの性質と程度を調査すること。
  • 長期的な研究を支援するための拡張可能でオープンソースのコードベースを提供すること。
  • 実世界の医療画像シナリオにおける既存の公平性低減手法の有効性を評価すること。

提案手法

  • X線、CT、MRI、超音波、OCT、眼底写真、皮膚画像など、2D、2.5D、3Dのモダリティをカバーする17の医療画像データセットを統合。
  • 一般用途(例:CLIP、ViT)および医療分野特化型(例:Med-PaLM、SAM-Med3D)の20の広く使われているFMsをサポート。
  • ゼロショット、線形プローブ、パラメータ効率の良いファインチューニング(PEFT)、プロンプトベースの適応、フルファインチューニングの複数の使用パラダイムを可能に。
  • モジュラー構成(データローダ、モデル、使用ラッパー、トレーナ、評価モジュール)を備えた、PyTorchに基づく統一されたコードベースを構築。
  • EqOdds、AUCΔ、ECEΔ、DSCΔ、予測整合性公平性など多様な公平性指標を適用。表現の公平性はt-SNEを用いて分析。
  • グループ再バランス、敵対的訓練、公平性制約など、最先端の公平性低減技術を統合。
Figure 1: Overview of the FairMedFM framework, a standardized pipeline to investigate fairness on diverse datasets (2D, 2.5D, and 3D), comprehensive functionalities (various FMs, tasks, usages, and debias algorithms), thorough evaluation metrics. The details are explained in Sec. 3 .
Figure 1: Overview of the FairMedFM framework, a standardized pipeline to investigate fairness on diverse datasets (2D, 2.5D, and 3D), comprehensive functionalities (various FMs, tasks, usages, and debias algorithms), thorough evaluation metrics. The details are explained in Sec. 3 .

実験結果

リサーチクエスチョン

  • RQ1異なる基盤モデルは、多様な医療画像データセットおよび感受性属性において、どのように公平性の面で性能を発揮するか?
  • RQ2さまざまなFMsの種別、使用戦略、下流タスクにおいて、有用性と公平性のトレードオフの性質と程度は何か?
  • RQ3既存の公平性低減手法は、医療画像分野のFMsにおけるバイアス低減にどの程度効果を発揮するか?
  • RQ4同じデータセットに適用した場合、異なるFMs間で公平性の乖離は一貫しているか?
  • RQ5FMsの潜在表現は公平性の結果とどのように関係しているか? また、表現空間の操作によってバイアス低減が可能か?

主な発見

  • CDダイアグラムによる統計的検定により、どのFMsも公平性または有用性において他と顕著な差がないことが確認され、モデル間で有意な性能差は認められなかった。
  • 性別、年齢、人種などの感受性属性にかかわらず、使用するFMsに関わらず顕著な公平性の不平等が存在し、モデル行動における恒常的バイアスが示された。
  • 有用性と公平性のトレードオフは顕著であり、FMsの種別や使用戦略によって変動し、最適なバランスを達成する単一のアプローチは存在しなかった。
  • 敵対的訓練やグループ再バランスなどの既存の公平性低減手法は、データセットやタスクをまたいでバイアス低減に限定的な有効性しか示さなかった。
  • 表現空間の分析から、データレベルでのグループ分離性(例:HAM10000)が高いほど公平性の不平等が顕著になることが判明し、バイアスの主要因としてデータ分布が関与している可能性が示唆された。
  • FairMedFMのオープンソースコードベースにより、拡張可能なベンチマークが可能となり、SAM-Med2D や FastSAM-3D といったプロンプトベースのモデルを用いたセグメンテーションを含む多様な利用事例をサポートした。
Figure 2: Bias in classification tasks. $\text{AUC}_{\Delta}$ is the fairness evaluation metric. “ $\dagger$ ” denotes vision-language models, and “ $*$ ” denotes pure vision models, where CLIP-ZS and CLIP-Adapt are not applicable.
Figure 2: Bias in classification tasks. $\text{AUC}_{\Delta}$ is the fairness evaluation metric. “ $\dagger$ ” denotes vision-language models, and “ $*$ ” denotes pure vision models, where CLIP-ZS and CLIP-Adapt are not applicable.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。