Skip to main content
QUICK REVIEW

[论文解读] Stability of clinical prediction models developed using statistical or machine learning methods

Richard D Riley, Gary S. Collins|arXiv (Cornell University)|Nov 2, 2022
Machine Learning in Healthcare参考文献 48被引用 8
一句话总结

本文提出一种框架,通过在自举样本上反复应用模型开发,利用不稳定性图、校准不稳定性曲线和不稳定性指数,评估临床环境中预测模型的不稳定性。结果表明,小样本数据会导致高不稳定性与校准偏差,提醒研究人员在验证或部署前评估模型的可靠性。

ABSTRACT

Clinical prediction models estimate an individual's risk of a particular health outcome, conditional on their values of multiple predictors. A developed model is a consequence of the development dataset and the chosen model building strategy, including the sample size, number of predictors and analysis method (e.g., regression or machine learning). Here, we raise the concern that many models are developed using small datasets that lead to instability in the model and its predictions (estimated risks). We define four levels of model stability in estimated risks moving from the overall mean to the individual level. Then, through simulation and case studies of statistical and machine learning approaches, we show instability in a model's estimated risks is often considerable, and ultimately manifests itself as miscalibration of predictions in new data. Therefore, we recommend researchers should always examine instability at the model development stage and propose instability plots and measures to do so. This entails repeating the model building steps (those used in the development of the original prediction model) in each of multiple (e.g., 1000) bootstrap samples, to produce multiple bootstrap models, and then deriving (i) a prediction instability plot of bootstrap model predictions (y-axis) versus original model predictions (x-axis), (ii) a calibration instability plot showing calibration curves for the bootstrap models in the original sample; and (iii) the instability index, which is the mean absolute difference between individuals' original and bootstrap model predictions. A case study is used to illustrate how these instability assessments help reassure (or not) whether model predictions are likely to be reliable (or not), whilst also informing a model's critical appraisal (risk of bias rating), fairness assessment and further validation requirements.

研究动机与目标

  • 解决由小样本数据开发的临床模型存在的预测不稳定性风险。
  • 识别模型不稳定性通常导致新数据中出现校准偏差。
  • 提出一种在模型开发过程中系统评估不稳定性的方法。
  • 通过不稳定性度量改进模型评估、公平性评估和验证规划。

提出的方法

  • 在1000个自举样本上重复模型开发,生成多个自举模型。
  • 创建预测不稳定性图,比较自举模型预测与原始模型预测。
  • 生成校准不稳定性图,展示在原始数据集中自举模型的校准曲线。
  • 将不稳定性指数计算为每个个体的原始预测与自举预测之间的平均绝对差。
  • 使用这些工具在模型开发过程中评估可靠性、偏差和公平性。
  • 在模拟研究和真实世界案例研究中应用该框架,以证明其有效性。

实验结果

研究问题

  • RQ1在临床预测模型中,模型不稳定性如何随不同样本大小和预测变量数量而变化?
  • RQ2当模型基于小样本数据开发时,不稳定性在多大程度上导致新数据中的校准偏差?
  • RQ3不稳定性图和不稳定性指数能否可靠地在外部验证前检测出不可靠的预测?
  • RQ4在小样本条件下,统计方法与机器学习方法在不稳定性敏感性方面有何差异?
  • RQ5不稳定性评估能否改善模型开发过程中的偏倚风险评估和公平性评估?

主要发现

  • 即使使用标准统计方法或机器学习方法,小样本数据中的模型不稳定性仍显著。
  • 不稳定性始终导致新数据中的校准偏差,从而损害模型的可靠性。
  • 不稳定性指数能有效量化自举样本中个体层面的预测变异性。
  • 预测不稳定性图揭示了原始预测与自举预测之间存在系统性偏离,未达到完美一致。
  • 校准不稳定性图凸显了自举模型在小样本中普遍存在校准性能差的问题。
  • 所提出的框架通过识别需要进一步验证或优化的不稳定模型,改善了模型评估。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。