Skip to main content
QUICK REVIEW

[論文レビュー] Prediction under interventions: evaluation of counterfactual performance using longitudinal observational data

Ruth H. Keogh, Nan van Geloven|arXiv (Cornell University)|Apr 19, 2023
Reliability and Agreement in Measurement被引用数 4
ひとこと要約

本稿は、時間発生イベントの結果を伴う縦断的観察データにおいて、介入的モデルの反事後的予測性能を評価するための新規手法を提示する。人工的な打ち切りと逆確率重み付けを組み合わせることで、仮想の治療戦略を模倣する合成された検証データセットを生成し、介入下でのコロナリティ、識別力(c-index、AUCt)、ブライアースコアの厳密な評価を可能にする。これは、臓器提供のような現実世界の状況における因果的予測の評価に向けた堅牢なフレームワークを提供する。

ABSTRACT

Predictions under interventions are estimates of what a person's risk of an outcome would be if they were to follow a particular treatment strategy, given their individual characteristics. Such predictions can give important input to medical decision making. However, evaluating predictive performance of interventional predictions is challenging. Standard ways of evaluating predictive performance do not apply when using observational data, because prediction under interventions involves obtaining predictions of the outcome under conditions that are different to those that are observed for a subset of individuals in the validation dataset. This work describes methods for evaluating counterfactual performance of predictions under interventions for time-to-event outcomes. This means we aim to assess how well predictions would match the validation data if all individuals had followed the treatment strategy under which predictions are made. We focus on counterfactual performance evaluation using longitudinal observational data, and under treatment strategies that involve sustaining a particular treatment regime over time. We introduce an estimation approach using artificial censoring and inverse probability weighting which involves creating a validation dataset that mimics the treatment strategy under which predictions are made. We extend measures of calibration, discrimination (c-index and cumulative/dynamic AUCt) and overall prediction error (Brier score) to allow assessment of counterfactual performance. The methods are evaluated using a simulation study, including scenarios in which the methods should detect poor performance. Applying our methods in the context of liver transplantation shows that our procedure allows quantification of the performance of predictions supporting crucial decisions on organ allocation.

研究の動機と目的

  • 仮想の治療介入下での結果を推定するモデルの予測性能を評価するうえで、重要なギャップを埋めること。
  • 時間発生イベントの結果を伴う縦断的観察データにおける反事後的性能の一般化可能なフレームワークを開発すること。
  • 観察された治療に条件づけられる選択バイアスのための欠陥を抱える既存の検証手法(例:サブセットアプローチ)の限界を克服すること。
  • 標準的な性能指標(コロナリティ、識別力(c-index、動的AUCt)、ブライアースコア)を、持続的治療戦略下の反事後的設定に拡張すること。
  • 肝移植のような現実の臨床的文脈における介入的予測の評価に実用的かつ妥当な手法を提供すること。

提案手法

  • 検証データセット内の個々の被験者に対して、人工的な打ち切りを施して仮想の治療戦略を模倣する:非介入の場合は時刻0で打ち切り、介入の場合はランドマーク時刻で打ち切り。
  • 選択バイアスを補正するために、治療および打ち切りの逆確率重み付け(IPACW)を適用する。介入戦略には時刻固定の重みを、非介入には時刻依存の重みを用いる。
  • 個人-ランドマークアプローチを用いて、特定の時点で移植を受けていたか否かを表す別個の検証データセット $V^1$ と $V^0$ を作成する。
  • 30日間隔ごとの区分的アプローチを用いて、時刻に依存するIPACW推定値を算出し、移植状態、待機リストからの除外、管理的打ち切りを考慮する。
  • 元の検証データ上でカプラン=マイヤー推定値から得られる追加の管理的打ち切り重み $G_c^{-1}(t)$ を組み込み、フォローアップ終了時のバイアスを是正する。
  • 重み付けされ、人工的に打ち切られた検証データセットに、標準的な性能指標(c-index、AUCt、ブライアースコア、コロナリティ)を適用し、反事後的性能を評価する。
Figure 1 : Directed acyclic graph (DAG) illustrating relationships between treatment $A$ , time-dependent covariates $L$ , baseline prognostic variables $P$ , and discrete time outcome $Y$ . The DAG is illustrated for a discrete-time setting where $Y_{k}=I(k-1\leq T<k)$ is an indicator of whether th
Figure 1 : Directed acyclic graph (DAG) illustrating relationships between treatment $A$ , time-dependent covariates $L$ , baseline prognostic variables $P$ , and discrete time outcome $Y$ . The DAG is illustrated for a discrete-time setting where $Y_{k}=I(k-1\leq T<k)$ is an indicator of whether th

実験結果

リサーチクエスチョン

  • RQ1観察データしか入手できない状況で、仮想の介入下での結果を推定するモデルの予測性能をどのように意味的に評価できるか?
  • RQ2時間発生イベントの文脈において、介入的予測を検証するのによく使われるサブセットアプローチにおける選択バイアスの影響は何か?
  • RQ3人工的な打ち切りと逆確率重み付けを組み合わせることで、持続的治療戦略下での有効な反事後的性能推定が可能になるか?
  • RQ4標準的な性能指標(コロナリティ、識別力、総合誤差)は、縦断的観察データにおける反事後的検証にどのように適合するか?
  • RQ5既知のシナリオ下で、本手法が既存のアプローチをどれほど上回って、不良なモデル性能を検出できるか?

主な発見

  • 提案手法は、人工的な打ち切りと逆確率重み付けを用いることで、仮想の介入下での偏りのない予測性能評価を可能にする有効な反事後的検証データセットを構築した。
  • 実務でよく使われるサブセットアプローチは、選択バイアスに起因し、性能評価が不正となることが示された。
  • 本手法は、c-index、動的AUCt、ブライアースコアといった主要な性能指標を、反事後的条件下で信頼性高く推定可能である。
  • 時刻に依存するIPACW重みは、非介入戦略において、移植や待機リストからの除外による打ち切りが時間とともに変化するため、不可欠であった。
  • 管理的打ち切り重み $G_c^{-1}(t)$ の組み込みは、特にフォローアップ期間が固定された研究では、性能推定の妥当性を維持するために不可欠であった。
  • 本手法は、肝移植データセットに成功裏に適用され、臓器提供のような高リスクの臨床意思決定文脈における実用性が示された。
Figure 2 : Simulation results: additive hazards model Scenario 1. Left panel: for the never treated strategy. Right panel: for the always treated strategy. Performance measures were obtained from the perfect validation data (black dots) and estimated from the observational validation data using the
Figure 2 : Simulation results: additive hazards model Scenario 1. Left panel: for the never treated strategy. Right panel: for the always treated strategy. Performance measures were obtained from the perfect validation data (black dots) and estimated from the observational validation data using the

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。