[Paper Review] Not Color Blind: AI Predicts Racial Identity from Black and White Retinal Vessel Segmentations
This study demonstrates that artificial intelligence can accurately predict racial identity from black-and-white retinal vessel segmentations—images previously assumed to lack racial information—using deep learning. Despite removing color and normalizing vessel brightness and width, convolutional neural networks achieved near-perfect race prediction (AUC-PR up to 0.995), revealing that racial bias can persist in AI models even when explicit pigmentation cues are removed.
Background: Artificial intelligence (AI) may demonstrate racial bias when skin or choroidal pigmentation is present in medical images. Recent studies have shown that convolutional neural networks (CNNs) can predict race from images that were not previously thought to contain race-specific features. We evaluate whether grayscale retinal vessel maps (RVMs) of patients screened for retinopathy of prematurity (ROP) contain race-specific features. Methods: 4095 retinal fundus images (RFIs) were collected from 245 Black and White infants. A U-Net generated RVMs from RFIs, which were subsequently thresholded, binarized, or skeletonized. To determine whether RVM differences between Black and White eyes were physiological, CNNs were trained to predict race from color RFIs, raw RVMs, and thresholded, binarized, or skeletonized RVMs. Area under the precision-recall curve (AUC-PR) was evaluated. Findings: CNNs predicted race from RFIs near perfectly (image-level AUC-PR: 0.999, subject-level AUC-PR: 1.000). Raw RVMs were almost as informative as color RFIs (image-level AUC-PR: 0.938, subject-level AUC-PR: 0.995). Ultimately, CNNs were able to detect whether RFIs or RVMs were from Black or White babies, regardless of whether images contained color, vessel segmentation brightness differences were nullified, or vessel segmentation widths were normalized. Interpretation: AI can detect race from grayscale RVMs that were not thought to contain racial information. Two potential explanations for these findings are that: retinal vessels physiologically differ between Black and White babies or the U-Net segments the retinal vasculature differently for various fundus pigmentations. Either way, the implications remain the same: AI algorithms have potential to demonstrate racial bias in practice, even when preliminary attempts to remove such information from the underlying images appear to be successful.
Motivation & Objective
- To investigate whether grayscale retinal vessel maps (RVMs) contain residual racial information not apparent in visual inspection.
- To assess whether AI models can detect race from RVMs even after normalization of brightness, width, and segmentation differences.
- To evaluate whether physiological differences in retinal vasculature or segmentation artifacts account for race prediction in RVMs.
- To highlight the risk of racial bias in AI-driven medical diagnostics, even when images are preprocessed to remove apparent racial markers.
- To demonstrate that race prediction from RVMs remains highly accurate despite attempts to eliminate color and pigmentation features.
Proposed method
- Retinal fundus images (RFIs) from 245 Black and White infants were collected, totaling 4,095 images.
- A U-Net model was used to generate retinal vessel maps (RVMs) from the RFIs, followed by thresholding, binarization, and skeletonization to create multiple RVM variants.
- Convolutional neural networks (CNNs) were trained to classify race using color RFIs, raw RVMs, and processed RVMs (thresholded, binarized, skeletonized).
- Performance was evaluated using area under the precision-recall curve (AUC-PR) at both image-level and subject-level.
- To test robustness, brightness differences and vessel width variations were normalized across groups before model training.
- The study compared model performance across image types to isolate whether race prediction relied on residual pigmentation or structural differences in vessel patterns.
Experimental results
Research questions
- RQ1Can AI models accurately predict racial identity from grayscale retinal vessel maps that were not previously considered to contain racial information?
- RQ2To what extent do residual differences in retinal vessel morphology or segmentation artifacts contribute to race prediction in RVMs?
- RQ3Does normalizing brightness and vessel width in RVMs eliminate the ability of AI to predict race?
- RQ4Are the observed race prediction capabilities due to physiological differences in retinal vasculature between Black and White infants, or due to model bias in segmentation?
- RQ5Can AI detect race from images that have been preprocessed to remove color and pigmentation cues?
Key findings
- CNNs achieved near-perfect race prediction from original color retinal fundus images, with an image-level AUC-PR of 0.999 and subject-level AUC-PR of 1.000.
- Raw retinal vessel maps (RVMs) alone enabled race prediction with high accuracy, achieving an image-level AUC-PR of 0.938 and subject-level AUC-PR of 0.995.
- Even after normalizing vessel brightness and width, CNNs maintained strong performance, indicating that race prediction was not dependent on these visual cues.
- The ability to predict race persisted across all RVM variants—thresholded, binarized, and skeletonized—demonstrating robustness to preprocessing.
- The findings suggest that either physiological differences in retinal vasculature or segmentation artifacts related to fundus pigmentation contribute to race detection in RVMs.
- The study reveals that AI models can inherit racial bias even when images are processed to remove apparent racial markers, posing a critical risk in clinical AI applications.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.