Skip to main content
QUICK REVIEW

[논문 리뷰] Prediction under interventions: evaluation of counterfactual performance using longitudinal observational data

Ruth H. Keogh, Nan van Geloven|arXiv (Cornell University)|2023. 04. 19.
Reliability and Agreement in Measurement인용 수 4
한 줄 요약

이 논문은 시간-이벤트 결과를 갖는 종단적 관찰 데이터에서 간섭 모델의 반사적 예측 성능을 평가하기 위한 새로운 방법을 제시한다. 인공적 케논과 역확률가중법을 결합하여 가상의 치료 전략을 모방하는 합성 검증 데이터셋을 생성함으로써, 치료에 대한 캘리브레이션, 분류능력(c-index, AUCt), 브리어 스코어의 엄밀한 평가가 가능해지며, 실생활 환경(예: 장기 이식)에서 인과적 예측을 평가하기 위한 강력한 프레임워크를 제공한다.

ABSTRACT

Predictions under interventions are estimates of what a person's risk of an outcome would be if they were to follow a particular treatment strategy, given their individual characteristics. Such predictions can give important input to medical decision making. However, evaluating predictive performance of interventional predictions is challenging. Standard ways of evaluating predictive performance do not apply when using observational data, because prediction under interventions involves obtaining predictions of the outcome under conditions that are different to those that are observed for a subset of individuals in the validation dataset. This work describes methods for evaluating counterfactual performance of predictions under interventions for time-to-event outcomes. This means we aim to assess how well predictions would match the validation data if all individuals had followed the treatment strategy under which predictions are made. We focus on counterfactual performance evaluation using longitudinal observational data, and under treatment strategies that involve sustaining a particular treatment regime over time. We introduce an estimation approach using artificial censoring and inverse probability weighting which involves creating a validation dataset that mimics the treatment strategy under which predictions are made. We extend measures of calibration, discrimination (c-index and cumulative/dynamic AUCt) and overall prediction error (Brier score) to allow assessment of counterfactual performance. The methods are evaluated using a simulation study, including scenarios in which the methods should detect poor performance. Applying our methods in the context of liver transplantation shows that our procedure allows quantification of the performance of predictions supporting crucial decisions on organ allocation.

연구 동기 및 목표

  • 관찰된 치료에 조건을 두어 발생하는 선택 편향로 인해 기존의 평가 방법이 취약한, 허구적 치료 간섭에 대한 결과를 추정하는 모델의 예측 성능 평가에 있어 근본적인 격차를 해소하고자 한다.
  • 시간-이벤트 결과를 갖는 종단적 관찰 데이터에서 반사적 성능을 평가할 수 있는 일반화 가능한 프레임워크를 개발하고자 한다.
  • 관측된 치료에 조건을 두어 발생하는 선택 편향 문제를 야기하는 기존의 검증 방법(예: 서브셋 접근법)의 한계를 극복하고자 한다.
  • 표준 성능 측정치인 캘리브레이션, 분류능력(c-index, 동적 AUCt), 브리어 스코어를 지속적인 치료 전략 하에서 반사적 설정으로 확장하고자 한다.
  • 실제 임상 환경(예: 간 이식)에서 간섭 예측을 평가하기 위한 실용적이고 검증된 방법을 제공하고자 한다.

제안 방법

  • 검증 데이터셋 내 개별 환자를 인공적으로 케논하여, 비간섭 전략을 시뮬레이션함: 비간섭의 경우 시간 0에서 케논, 간섭의 경우 랜드마크 시간에서 케논.
  • 선택 편향을 보정하기 위해 치료 및 케논의 역확률가중법(IPACW)을 적용함. 간섭 전략에는 시간 고정 가중치를, 비간섭에는 시간에 따라 변하는 가중치를 사용함.
  • 개별 환자-랜드마크 접근법을 사용하여, 특정 시점에 간이 이식을 받은 경우와 받지 않은 경우를 각각 나타내는 별도의 검증 데이터셋 $V^1$ 및 $V^0$를 생성함.
  • 30일 간격의 조각별 접근법을 사용해 시간에 따라 변하는 IPACW를 추정함. 이는 이식 상태, 대기열에서 제거된 상태, 관리적 케논을 고려함.
  • 원본 검증 데이터에서 케파만-메이어 추정치를 기반으로 유도된 추가적인 관리적 케논 가중치 $G_c^{-1}(t)$를 포함하여, 추적 종료 시점의 편향을 보정함.
  • 가중치가 부여되고 인공적으로 케논된 검증 데이터셋에 표준 성능 측정치(c-index, AUCt, 브리어 스코어, 캘리브레이션)를 적용하여 반사적 성능을 평가함.
Figure 1 : Directed acyclic graph (DAG) illustrating relationships between treatment $A$ , time-dependent covariates $L$ , baseline prognostic variables $P$ , and discrete time outcome $Y$ . The DAG is illustrated for a discrete-time setting where $Y_{k}=I(k-1\leq T<k)$ is an indicator of whether th
Figure 1 : Directed acyclic graph (DAG) illustrating relationships between treatment $A$ , time-dependent covariates $L$ , baseline prognostic variables $P$ , and discrete time outcome $Y$ . The DAG is illustrated for a discrete-time setting where $Y_{k}=I(k-1\leq T<k)$ is an indicator of whether th

실험 결과

연구 질문

  • RQ1관찰 데이터만 제공되는 상황에서, 허구적 간섭에 대한 결과를 추정하는 모델의 예측 성능을 어떻게 의미 있게 평가할 수 있는가?
  • RQ2시간-이벤트 설정에서 간섭 예측을 검증하는 데 일반적으로 사용되는 서브셋 접근법에서 선택 편향이 미치는 영향은 무엇인가?
  • RQ3인공적 케논과 역확률가중법을 조합하면 지속적인 치료 전략 하에서 유효한 반사적 성능 추정치를 도출할 수 있는가?
  • RQ4표준 성능 측정치인 캘리브레이션, 분류능력, 총 오차는 종단적 관찰 데이터에서 반사적 검증에 어떻게 적응하는가?
  • RQ5기존의 접근법에 비해 제안된 방법이 알려진 상황에서 열악한 모델 성능을 탐지하는 데 얼마나 뛰어나게 성능을 발휘하는가?

주요 결과

  • 제안된 방법은 인공적 케논과 역확률가중법을 통해 유효한 반사적 검증 데이터셋을 성공적으로 생성하여, 허구적 간섭 하에서의 예측 성능 평가를 편향 없이 가능하게 하였다.
  • 실제로 널리 사용되는 서브셋 접근법은 선택 편향에 취약하여 잘못된 성능 평가를 초래하는 것으로 밝혀졌다.
  • 이 방법은 반사적 조건 하에서 c-index, 동적 AUCt, 브리어 스코어와 같은 핵심 성능 지표를 신뢰성 있게 추정할 수 있도록 하였다.
  • 비간섭 전략에서는 이식 및 대기열 제거로 인한 케논이 시간에 따라 변하기 때문에, 시간에 따라 변하는 IPACW 가중치가 필수적이었다.
  • 관리적 케논 가중치 $G_c^{-1}(t)$의 포함은 특히 고정된 추적 기간을 갖는 연구에서 성능 추정의 타당성을 유지하는 데 핵심적이었다.
  • 이 방법은 간 이식 데이터셋에 성공적으로 적용되어, 장기 이식과 같은 고위험 임상 의사결정 환경에서의 유용성을 입증하였다.
Figure 2 : Simulation results: additive hazards model Scenario 1. Left panel: for the never treated strategy. Right panel: for the always treated strategy. Performance measures were obtained from the perfect validation data (black dots) and estimated from the observational validation data using the
Figure 2 : Simulation results: additive hazards model Scenario 1. Left panel: for the never treated strategy. Right panel: for the always treated strategy. Performance measures were obtained from the perfect validation data (black dots) and estimated from the observational validation data using the

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.