[论文解读] Risk of Bias in Chest Radiography Deep Learning Foundation Models
本研究利用CheXpert数据集(n=127,118张影像)评估了胸部X光摄影深度学习基础模型在生物性别和种族之间的偏见。研究发现显著的性能差异——女性在'无异常'标签上的准确率较低,非裔患者在'胸腔积液'标签上的准确率也较低,表明存在系统性偏见,威胁临床安全与公平性。
Purpose: To analyze a recently published chest radiography foundation model for the presence of biases that could lead to subgroup performance disparities across biological sex and race. Materials and Methods: This retrospective study used 127,118 chest radiographs from 42,884 patients (mean age, 63 [SD] 17 years; 23,623 male, 19,261 female) from the CheXpert dataset collected between October 2002 and July 2017. To determine the presence of bias in features generated by a chest radiography foundation model and baseline deep learning model, dimensionality reduction methods together with two-sample Kolmogorov-Smirnov tests were used to detect distribution shifts across sex and race. A comprehensive disease detection performance analysis was then performed to associate any biases in the features to specific disparities in classification performance across patient subgroups. Results: Ten out of twelve pairwise comparisons across biological sex and race showed statistically significant differences in the studied foundation model, compared with four significant tests in the baseline model. Significant differences were found between male and female (P < .001) and Asian and Black patients (P < .001) in the feature projections that primarily capture disease. Compared with average model performance across all subgroups, classification performance on the 'no finding' label dropped between 6.8% and 7.8% for female patients, and performance in detecting 'pleural effusion' dropped between 10.7% and 11.6% for Black patients. Conclusion: The studied chest radiography foundation model demonstrated racial and sex-related bias leading to disparate performance across patient subgroups and may be unsafe for clinical applications.
研究动机与目标
- 调查胸部X光摄影基础模型是否在生物性别和种族之间表现出偏见。
- 评估模型中特征分布是否在不同子群体间存在显著差异,以识别潜在偏见。
- 评估此类偏见是否转化为疾病分类性能中的可测量差异。
- 比较基础模型与基线深度学习模型之间的偏见水平。
- 确定观察到的性能差异在真实世界AI应用中的临床安全性影响。
提出的方法
- 对2002–2017年间来自CheXpert数据集的127,118张胸部X光片进行回顾性分析,包含性别和种族的标注信息。
- 应用降维技术提取并比较不同性别和种族子群体的特征表示。
- 使用两样本Kolmogorov-Smirnov检验检测模型学习特征在不同子群体间分布变化的统计显著性。
- 在不同性别和种族子群体中评估基础模型在关键诊断标签(如'无异常'、'胸腔积液')上的性能。
- 比较基础模型与基线深度学习模型之间的偏见指标和性能差异。
- 对分类性能差异进行定量分析,重点关注表现较差子群体的准确率绝对下降值。
实验结果
研究问题
- RQ1基础模型中,男性与女性患者在特征分布上是否存在统计显著差异?
- RQ2不同种族子群体(如亚裔与非裔)在模型潜在空间中的特征表示是否表现出差异?
- RQ3基础模型在不同性别和种族子群体中的疾病分类性能是否存在可测量的差异?
- RQ4基础模型中的偏见水平与基线深度学习模型相比如何?
- RQ5观察到的特征层面偏见在多大程度上与临床相关的性能差异相关?
主要发现
- 在性别和种族之间的12组两两比较中,基础模型的特征表现出10组统计显著的分布偏移,而基线模型仅4组。
- 在捕捉与疾病相关模式的特征投影中,男性与女性患者(p < .001)以及亚裔与非裔患者(p < .001)之间均发现显著差异。
- 与所有子群体的平均值相比,女性患者在'无异常'标签上的分类性能下降了6.8%至7.8%。
- 与整体平均水平相比,非裔患者在检测'胸腔积液'时的性能下降了10.7%至11.6%。
- 基础模型表现出具有临床相关性的性能差异,表明学习特征中的偏见确实转化为现实世界中的性能差距。
- 本研究结论认为,由于存在与种族和性别相关的性能差异,该模型不适合临床部署。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。