[論文レビュー] Risk of Bias in Chest Radiography Deep Learning Foundation Models
本研究では、チェクサントデータセット(n=127,118スキャン)を用いて、生物学的性別および人種に跨るバイアスを、胸部レントゲン撮影用のディーブラーニング基盤モデルに対して評価した。特に『異常なし』の分類で女性において、および『胸膜効果』の分類でブラック・アフリカ系患者において顕著な性能格差が認められ、臨床的セーフティおよび公平性を損なうシステム的バイアスが存在することが示された。
Purpose: To analyze a recently published chest radiography foundation model for the presence of biases that could lead to subgroup performance disparities across biological sex and race. Materials and Methods: This retrospective study used 127,118 chest radiographs from 42,884 patients (mean age, 63 [SD] 17 years; 23,623 male, 19,261 female) from the CheXpert dataset collected between October 2002 and July 2017. To determine the presence of bias in features generated by a chest radiography foundation model and baseline deep learning model, dimensionality reduction methods together with two-sample Kolmogorov-Smirnov tests were used to detect distribution shifts across sex and race. A comprehensive disease detection performance analysis was then performed to associate any biases in the features to specific disparities in classification performance across patient subgroups. Results: Ten out of twelve pairwise comparisons across biological sex and race showed statistically significant differences in the studied foundation model, compared with four significant tests in the baseline model. Significant differences were found between male and female (P < .001) and Asian and Black patients (P < .001) in the feature projections that primarily capture disease. Compared with average model performance across all subgroups, classification performance on the 'no finding' label dropped between 6.8% and 7.8% for female patients, and performance in detecting 'pleural effusion' dropped between 10.7% and 11.6% for Black patients. Conclusion: The studied chest radiography foundation model demonstrated racial and sex-related bias leading to disparate performance across patient subgroups and may be unsafe for clinical applications.
研究の動機と目的
- 胸部レントゲン用基盤モデルが生物学的性別および人種に跨るバイアスを示すかどうかを調査すること。
- モデル内の特徴分布がサブグループ間で顕著に異なるかどうかを評価すること。
- このようなバイアスが、疾患分類性能における測定可能な格差にどのように反映されるかを評価すること。
- 基盤モデルとベースラインディーブラーニングモデルとの間でのバイアスレベルを比較すること。
- 観察された性能格差が、実世界のAIアプリケーションにおける臨床的セーフティに及ぼす影響を特定すること。
提案手法
- 2002年から2017年までの間のチェクサントデータセットから抽出された127,118枚の胸部レントゲン画像の後向き分析(性別および人種のラベルを付与済み)。
- 次元削減技術を用いて、性別および人種サブグループ間での特徴表現を抽出および比較した。
- 2標本Kolmogorov-Smirnov検定を用いて、サブグループ間での学習済み特徴の分布シフトが統計的に有意かどうかを同定した。
- 基盤モデルの、『異常なし』や『胸膜効果』などの主要な診断ラベルにおける、性別および人種サブグループごとの性能評価。
- 基盤モデルとベースラインディーブラーニングモデルとの間でのバイアス指標および性能格差の比較。
- 特に性能が低いサブグループにおける精度の絶対的低下を焦点に、分類性能差の定量的分析。
実験結果
リサーチクエスチョン
- RQ1基盤モデルにおいて、男性と女性の患者間で特徴分布に統計的に有意な差が認められるか?
- RQ2アジア系とブラック・アフリカ系などの人種サブグループ間で、モデルの潜在空間における特徴表現が異なるか?
- RQ3基盤モデルが、性別および人種サブグループに跨る疾患分類性能に測定可能な格差を示すか?
- RQ4基盤モデルのバイアスレベルは、ベースラインディーブラーニングモデルと比較してどの程度異なるか?
- RQ5観察された特徴レベルのバイアスは、臨床的に重要な性能格差とどの程度相関しているか?
主な発見
- 性別および人種に跨る12組のペアワイズ比較のうち、基盤モデルでは10組で統計的に有意な特徴分布シフトが認められたのに対し、ベースラインモデルでは4組であった。
- 疾患関連パターンを捉える特徴の投影において、男性と女性の患者間(p < .001)、およびアジア系とブラック・アフリカ系の患者間(p < .001)で顕著な差が認められた。
- 女性患者では、『異常なし』ラベルの分類性能が、全サブグループの平均と比較して6.8%〜7.8%低下した。
- ブラック・アフリカ系の患者では、『胸膜効果』の検出性能が、全平均と比較して10.7%〜11.6%低下した。
- 基盤モデルは臨床的に意味のある格差を示しており、学習済み特徴に内在するバイアスが、実世界の性能ギャップに反映されていることを示した。
- 本研究は、人種および性別に起因する性能格差のため、モデルは臨床的展開には安全でないとの結論に至った。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。