[论文解读] Coarse race data conceals disparities in clinical risk score performance
本论文表明,粒度化的种族数据揭示了临床风险评分表现中的显著差异,这些差异在使用粗略种族类别时被隐藏;基于418K次急诊就诊,覆盖26个粒度分组。
Healthcare data in the United States often records only a patient's coarse race group: for example, both Indian and Chinese patients are typically coded as "Asian." It is unknown, however, whether this coarse coding conceals meaningful disparities in the performance of clinical risk scores across granular race groups. Here we show that it does. Using data from 418K emergency department visits, we assess clinical risk score performance disparities across 26 granular groups for three outcomes, five risk scores, and four performance metrics. Across outcomes and metrics, we show that the risk scores exhibit significant granular performance disparities within coarse race groups. In fact, variation in performance within coarse groups often *exceeds* the variation between coarse groups. We explore why these disparities arise, finding that outcome rates, feature distributions, and the relationships between features and outcomes all vary significantly across granular groups. Our results suggest that healthcare providers, hospital systems, and machine learning researchers should strive to collect, release, and use granular race data in place of coarse race data, and that existing analyses may significantly underestimate racial disparities in performance.
研究动机与目标
- 在医疗分析中因粗略种族分组内存在异质性而需要粒度化的种族数据的动机。
- 量化粒度化种族子组之间预测性风险评分的差异。
- 评估粗略种族分析在多大程度上低估风险评分表现中的种族差异。
- 研究驱动粒度差异的数据分布因素(结果频率、特征分布以及 X→y 关系)。
提出的方法
- 使用来自 BIDMC 的 418K 次急诊就诊,数据包含自我认定的粗粒度和粒度化种族类别的 MIMIC-IV-ED 数据。
- 在三个急诊结局上评估五个风险评分(两个临床评分和三个 ML 模型)。
- 对粗粒度组和粒度组计算四个性能指标(AUPRC、AUROC、FPR、FNR),并给出 95% 置信区间。
- 在将粒度子组与粗组比较时,对多重假设检验应用 Bonferroni 校正。
- 通过样本量检查、结果频率、特征分布以及 p(y|X) 的变异来分析数据分布对差异的贡献。
- 使用替代模型(LR 和 XGBoost)重复 ML 结果以验证稳健性。
实验结果
研究问题
- RQ1在同一粗粒度种族类别内,粒度化种族分组的预测性风险评分表现是否存在显著差异?
- RQ2在同一粗粒度组内的表现变异是否大于不同粗粒度组之间的变异,表明粗组分析隐藏差异?
- RQ3哪些数据分布因素(样本量、结果频率、特征分布、特征-结果关系)驱动粒度差异对性能的影响?
- RQ4是否应收集并使用粒度化种族数据来更好地评估和解决临床风险评分中的公平性问题?
主要发现
- 粒度化种族分组在多种结局和指标上显示出显著的表现差异,而粗粒度组未能捕捉到这些差异。
- 同一粗粒度组内的性能变异性常常与粗粒度组之间的变异性相当甚至更大,有时超过两倍以上。
- 结局频率在粒度组之间有显著差异,导致 AUPRC、FPR、FNR 等指标的差异。
- 特征分布和特征-结果关系在粒度组之间存在差异,表明协变量分布漂移和不同组之间的预测信号差异。
- 回归分析表明粒度种族交互显著改善了结局的拟合,暗示 p(y|X) 在粗组内的粒度组之间存在差异。
- 分诊急性和特定合并症具有组特异的预测重要性,提示风险评分在不同粒度组之间可能需要不同的校准。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。