Skip to main content
QUICK REVIEW

[论文解读] When accurate prediction models yield harmful self-fulfilling prophecies

Wouter A. C. van Amsterdam, Nan van Geloven|arXiv (Cornell University)|Dec 2, 2023
Explainable Artificial Intelligence (XAI)被引用 6
一句话总结

本文表明,即使在部署后仍保持强大区分度和校准性,准确的医疗预测模型在用于指导临床决策时,也可能产生有害的自我实现预言。作者形式化了此类模型因强化现有偏见而对患者亚群体造成系统性伤害的条件,表明仅靠预测准确性不足以确保临床决策支持系统中的安全部署。

ABSTRACT

Prediction models are popular in medical research and practice. By predicting an outcome of interest for specific patients, these models may help inform difficult treatment decisions, and are often hailed as the poster children for personalized, data-driven healthcare. We show however, that using prediction models for decision making can lead to harmful decisions, even when the predictions exhibit good discrimination after deployment. These models are harmful self-fulfilling prophecies: their deployment harms a group of patients but the worse outcome of these patients does not invalidate the predictive power of the model. Our main result is a formal characterization of a set of such prediction models. Next we show that models that are well calibrated before and after deployment are useless for decision making as they made no change in the data distribution. These results point to the need to revise standard practices for validation, deployment and evaluation of prediction models that are used in medical decisions.

研究动机与目标

  • 调查验证研究中的预测准确性是否足以确保结果预测模型在临床决策支持中安全部署。
  • 识别准确预测模型在医疗决策中导致有害自我实现预言的条件。
  • 形式化模型部署如何以保持模型性能但恶化特定亚群体护理的方式改变患者结局的机制。
  • 挑战医学中仅依赖区分度和校准性进行模型验证与部署的常规做法。

提出的方法

  • 作者构建了一个因果框架,使用潜在结果和干预分布来建模预测模型、治疗政策与患者结局之间的相互作用。
  • 定义自我实现预言为一种情形:模型的部署导致某一群体结局变差,但模型在部署后仍保持良好的预测性能。
  • 该方法包括模型在部署前和部署后校准性的正式条件,表明在两种分布下均校准良好的模型对决策无效。
  • 关键方程刻画了部署前后结局期望值的差异,识别出模型驱动政策导致伤害的时机。
  • 分析采用二元协变量设定,证明即使部署后区分度很强,仍存在非平凡的模型子集会导致有害结局。
  • 该框架指出,当治疗政策根据预测结果改变,且结局随之与预测结果一致地变化时,伤害即发生,从而强化偏见。
Figure 1: Some outcome prediction models yield harmful self-fulfilling prophecies when used to guide treatment decisions, meaning the new policy harms a subgroup of patients but the prediction model has good discrimination post-deployment.
Figure 1: Some outcome prediction models yield harmful self-fulfilling prophecies when used to guide treatment decisions, meaning the new policy harms a subgroup of patients but the prediction model has good discrimination post-deployment.

实验结果

研究问题

  • RQ1在何种条件下,部署准确的预测模型会导致患者亚群体结局更差,尽管其预测性能在部署后仍良好?
  • RQ2为何基于准确模型的有害政策无法被标准评估指标(如区分度和校准性)检测到?
  • RQ3在何种情况下,预测模型在部署前后均表现良好校准性?这对临床决策意味着什么?
  • RQ4为何模型在验证中表现准确,却仍可能持续或加剧现有医疗不平等?
  • RQ5预测模型在现实临床部署中既准确又无害,需满足哪些条件?

主要发现

  • 存在非平凡的准确预测模型子集,在临床环境中产生有害的自我实现预言,即部署后导致特定患者亚群体结局更差,尽管部署后仍具备强大区分度。
  • 在部署前后均表现良好校准性的模型对决策无效,因为这表明在新政策下数据分布未发生改变。
  • 当治疗政策基于预测结果制定,且结局变化与模型预测结果一致时,伤害即发生,即使模型仍保持准确。
  • 部署前后结局期望值的差异取决于治疗政策变动与有治疗与无治疗下结局差异的交互作用。
  • 当且仅当治疗政策保持不变,或各亚群体在有治疗与无治疗下的结局概率相等时,模型在部署分布上才表现校准。
  • 本文表明,仅靠预测准确性无法确保安全部署,因为模型可能准确但持续或加剧健康不平等。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。