[Paper Review] Risk of Bias in Chest Radiography Deep Learning Foundation Models
This study evaluates a chest radiography deep learning foundation model for bias across biological sex and race using the CheXpert dataset (n=127,118 scans). It identifies significant performance disparities—particularly lower accuracy for females on 'no finding' and for Black patients on 'pleural effusion'—indicating systemic bias that undermines clinical safety and fairness.
Purpose: To analyze a recently published chest radiography foundation model for the presence of biases that could lead to subgroup performance disparities across biological sex and race. Materials and Methods: This retrospective study used 127,118 chest radiographs from 42,884 patients (mean age, 63 [SD] 17 years; 23,623 male, 19,261 female) from the CheXpert dataset collected between October 2002 and July 2017. To determine the presence of bias in features generated by a chest radiography foundation model and baseline deep learning model, dimensionality reduction methods together with two-sample Kolmogorov-Smirnov tests were used to detect distribution shifts across sex and race. A comprehensive disease detection performance analysis was then performed to associate any biases in the features to specific disparities in classification performance across patient subgroups. Results: Ten out of twelve pairwise comparisons across biological sex and race showed statistically significant differences in the studied foundation model, compared with four significant tests in the baseline model. Significant differences were found between male and female (P < .001) and Asian and Black patients (P < .001) in the feature projections that primarily capture disease. Compared with average model performance across all subgroups, classification performance on the 'no finding' label dropped between 6.8% and 7.8% for female patients, and performance in detecting 'pleural effusion' dropped between 10.7% and 11.6% for Black patients. Conclusion: The studied chest radiography foundation model demonstrated racial and sex-related bias leading to disparate performance across patient subgroups and may be unsafe for clinical applications.
Motivation & Objective
- To investigate whether chest radiography foundation models exhibit bias across biological sex and race.
- To assess if feature distributions from the model differ significantly between subgroups, indicating potential bias.
- To evaluate whether such biases translate into measurable disparities in disease classification performance.
- To compare bias levels between a foundation model and a baseline deep learning model.
- To determine the clinical safety implications of observed performance disparities in real-world AI applications.
Proposed method
- Retrospective analysis of 127,118 chest radiographs from the CheXpert dataset (2002–2017), with demographic labels for sex and race.
- Application of dimensionality reduction techniques to extract and compare feature representations across sex and racial subgroups.
- Use of two-sample Kolmogorov-Smirnov tests to detect statistically significant distribution shifts in learned features between subgroups.
- Performance evaluation of the foundation model on key diagnostic labels (e.g., 'no finding', 'pleural effusion') across sex and race subgroups.
- Comparison of bias metrics and performance disparities between the foundation model and a baseline deep learning model.
- Quantitative analysis of classification performance differences, focusing on absolute drops in accuracy for underperforming subgroups.
Experimental results
Research questions
- RQ1Are there statistically significant differences in feature distributions between male and female patients in the foundation model?
- RQ2Do racial subgroups (e.g., Asian vs. Black) exhibit divergent feature representations in the model's latent space?
- RQ3Does the foundation model show measurable performance disparities in disease classification across sex and racial subgroups?
- RQ4How do bias levels in the foundation model compare to those in a baseline deep learning model?
- RQ5To what extent do observed feature-level biases correlate with clinically relevant performance disparities?
Key findings
- Ten out of twelve pairwise comparisons across sex and race showed statistically significant distribution shifts in the foundation model’s features, compared to four in the baseline model.
- Significant differences were found between male and female patients (p < .001) and between Asian and Black patients (p < .001) in feature projections capturing disease-related patterns.
- For female patients, classification performance on the 'no finding' label dropped by 6.8% to 7.8% compared to the average across all subgroups.
- For Black patients, performance in detecting 'pleural effusion' dropped by 10.7% to 11.6% relative to the overall average.
- The foundation model demonstrated clinically relevant disparities, indicating that bias in learned features translates into real-world performance gaps.
- The study concludes that the model is unsafe for clinical deployment due to these race- and sex-related performance disparities.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.