Skip to main content
QUICK REVIEW

[論文レビュー] Stability of clinical prediction models developed using statistical or machine learning methods

Richard D Riley, Gary S. Collins|arXiv (Cornell University)|Nov 2, 2022
Machine Learning in Healthcare参考文献 48被引用数 8
ひとこと要約

本論文は、ブートストラップ標本に対して繰り返しモデル開発を適用することで、臨床的設定における予測モデルの不安定性を評価するフレームワークを提案する。不安定性プロット、キャリブレーション不安定性曲線、および不安定性インデックスを用いる。小さなデータセットでは高い不安定性とキャリブレーションのずれが生じることを示し、検証や導入の前に信頼性を評価する必要があることを強調する。

ABSTRACT

Clinical prediction models estimate an individual's risk of a particular health outcome, conditional on their values of multiple predictors. A developed model is a consequence of the development dataset and the chosen model building strategy, including the sample size, number of predictors and analysis method (e.g., regression or machine learning). Here, we raise the concern that many models are developed using small datasets that lead to instability in the model and its predictions (estimated risks). We define four levels of model stability in estimated risks moving from the overall mean to the individual level. Then, through simulation and case studies of statistical and machine learning approaches, we show instability in a model's estimated risks is often considerable, and ultimately manifests itself as miscalibration of predictions in new data. Therefore, we recommend researchers should always examine instability at the model development stage and propose instability plots and measures to do so. This entails repeating the model building steps (those used in the development of the original prediction model) in each of multiple (e.g., 1000) bootstrap samples, to produce multiple bootstrap models, and then deriving (i) a prediction instability plot of bootstrap model predictions (y-axis) versus original model predictions (x-axis), (ii) a calibration instability plot showing calibration curves for the bootstrap models in the original sample; and (iii) the instability index, which is the mean absolute difference between individuals' original and bootstrap model predictions. A case study is used to illustrate how these instability assessments help reassure (or not) whether model predictions are likely to be reliable (or not), whilst also informing a model's critical appraisal (risk of bias rating), fairness assessment and further validation requirements.

研究の動機と目的

  • 小さなデータセットから開発された臨床モデルにおける不安定な予測のリスクに対処すること。
  • モデルの不安定性が新しいデータにおいてしばしばキャリブレーションのずれを引き起こすことを特定すること。
  • モデル開発段階で不安定性を評価するための体系的な手法を提案すること。
  • 不安定性指標を通じて、モデルの評価、公平性の評価、および検証計画の改善を図ること。

提案手法

  • 1000個のブートストラップ標本に対してモデル開発を繰り返し、複数のブートストラップモデルを生成する。
  • オリジナルモデルの予測と比較して、ブートストラップモデルの予測を示す予測不安定性プロットを作成する。
  • オリジナルデータセットにおけるブートストラップモデルのキャリブレーション曲線を示すキャリブレーション不安定性プロットを生成する。
  • 個々の被験者について、オリジナル予測とブートストラップ予測の絶対差の平均として、不安定性インデックスを計算する。
  • これらのツールを用いて、モデル開発段階で信頼性、バイアス、公平性を評価する。
  • シミュレーション研究および実世界の事例研究にこのフレームワークを適用し、その有効性を示す。

実験結果

リサーチクエスチョン

  • RQ1臨床予測モデルにおいて、さまざまなサンプルサイズと予測子数の下で、モデルの不安定性はどのように変化するか?
  • RQ2小さなデータセットからモデルを構築した場合、不安定性が新しいデータにおけるキャリブレーションのずれにどの程度寄与するか?
  • RQ3不安定性プロットおよび不安定性インデックスは、外部検証の前段階で信頼性の低い予測を信頼性を持って検出できるか?
  • RQ4統計的手法と機械学習手法は、小さな標本条件下で不安定性に対してどのように異なる感受性を示すか?
  • RQ5不安定性の評価は、モデル開発におけるバイアスのリスク評価および公平性の評価を改善できるか?

主な発見

  • 標準的な統計的手法や機械学習手法を用いても、小さなデータセットではモデルの不安定性が顕著に現れる。
  • 不安定性は常に新しいデータにおけるキャリブレーションのずれを引き起こし、モデルの信頼性を損なう。
  • 不安定性インデックスは、ブートストラップ標本間での個別予測のばらつきを効果的に定量化できる。
  • 予測不安定性プロットは、オリジナル予測とブートストラップ予測の間で完全な一致から系統的なずれを明らかにする。
  • キャリブレーション不安定性プロットは、特に小さな標本において、ブートストラップモデル全体でのキャリブレーション性能の悪さを強調する。
  • 提案されたフレームワークにより、信頼性の低いモデルを特定し、さらなる検証や精錬を要するモデルを特定することで、モデルの評価が向上する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。