Skip to main content
QUICK REVIEW

[Paper Review] Coarse race data conceals disparities in clinical risk score performance

Rajiv Movva, Divya Shanmugam|arXiv (Cornell University)|Apr 18, 2023
Emergency and Acute Care Studies11 citations
TL;DR

The paper demonstrates that granular race data reveals significant disparities in clinical risk-score performance that are hidden when using coarse race categories, using 418K emergency department visits across 26 granular groups.

ABSTRACT

Healthcare data in the United States often records only a patient's coarse race group: for example, both Indian and Chinese patients are typically coded as "Asian." It is unknown, however, whether this coarse coding conceals meaningful disparities in the performance of clinical risk scores across granular race groups. Here we show that it does. Using data from 418K emergency department visits, we assess clinical risk score performance disparities across 26 granular groups for three outcomes, five risk scores, and four performance metrics. Across outcomes and metrics, we show that the risk scores exhibit significant granular performance disparities within coarse race groups. In fact, variation in performance within coarse groups often *exceeds* the variation between coarse groups. We explore why these disparities arise, finding that outcome rates, feature distributions, and the relationships between features and outcomes all vary significantly across granular groups. Our results suggest that healthcare providers, hospital systems, and machine learning researchers should strive to collect, release, and use granular race data in place of coarse race data, and that existing analyses may significantly underestimate racial disparities in performance.

Motivation & Objective

  • Motivate the need for granular race data in healthcare analytics due to heterogeneity within coarse race groups.
  • Quantify how predictive risk scores differ across granular race subgroups.
  • Assess the extent to which coarse-race analyses underestimate racial disparities in risk-score performance.
  • Investigate data-distribution factors (outcome rates, feature distributions, and X→y relationships) that drive granular disparities.

Proposed method

  • Use MIMIC-IV-ED data with self-identified coarse and granular race categories for 418K ED visits from BIDMC.
  • Evaluate five risk scores (two clinical scores and three ML models) on three ED outcomes.
  • Compute four performance metrics (AUPRC, AUROC, FPR, FNR) across coarse and granular groups with 95% CIs.
  • Apply Bonferroni correction for multiple hypothesis testing when comparing granular subgroups to coarse groups.
  • Analyze the contribution of data distributions to disparities via sample size checks, outcome frequencies, feature distributions, and p(y|X) variation.
  • Replicate ML results using alternative models (LR and XGBoost) to verify robustness.

Experimental results

Research questions

  • RQ1Do predictive risk-score performances differ significantly across granular race groups within the same coarse race category?
  • RQ2Is within-coarse-group variation in performance larger than between-coarse-group variation, indicating hidden disparities under coarse-group analyses?
  • RQ3What data-distribution factors (sample size, outcome frequencies, feature distributions, feature-outcome relationships) drive granular disparities in performance?
  • RQ4Should granular race data be collected and used to better assess and address fairness in clinical risk scores?

Key findings

  • Granular race groups exhibit significant performance disparities not captured by coarse race groups across multiple outcomes and metrics.
  • Within-coarse-group variation in performance is often comparable to or larger than between-coarse-group variation, sometimes exceeding it by more than twofold.
  • Outcome frequencies differ meaningfully across granular groups, contributing to disparities in metrics like AUPRC, FPR, and FNR.
  • Feature distributions and feature–outcome relationships vary across granular groups, indicating covariate shift and differing predictive signals across groups.
  • Regression analyses show granular-race interactions significantly improve fit for predicting outcomes, implying p(y|X) varies across granular groups within coarse groups.
  • Triage acuity and specific comorbidities have group-specific predictive importance, suggesting that risk scores may be differently calibrated across granular groups.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.