Skip to main content
QUICK REVIEW

[논문 리뷰] Collaborative-controlled LASSO for Constructing Propensity Score-based Estimators in High-Dimensional Data

Cheng Ju, Richard Wyss|arXiv (Cornell University)|2017. 06. 30.
Advanced Causal Inference Techniques참고 문헌 28인용 수 9
한 줄 요약

이 논문은 고차원 데이터에서의 성향스코어 추정을 위해 협업 제어 LASSO 방법을 도입하며, C-TMLE를 사용하여 치료 예측과 인과적 효과 추정을 동시에 최적화한다. C-TMLE 기반 모델 선택이 전통적인 외부 교차검증보다 편향을 줄이고 평균 치료 효과 추정의 신뢰구간 커버리지 향상에 뛰어난 성능을 보임을 입증한다.

ABSTRACT

Propensity score (PS) based estimators are increasingly used for causal inference in observational studies. However, model selection for PS estimation in high-dimensional data has received little attention. In these settings, PS models have traditionally been selected based on the goodness-of-fit for the treatment mechanism itself, without consideration of the causal parameter of interest. Collaborative minimum loss-based estimation (C-TMLE) is a novel methodology for causal inference that takes into account information on the causal parameter of interest when selecting a PS model. This "collaborative learning" considers variable associations with both treatment and outcome when selecting a PS model in order to minimize a bias-variance trade off in the estimated treatment effect. In this study, we introduce a novel approach for collaborative model selection when using the LASSO estimator for PS estimation in high-dimensional covariate settings. To demonstrate the importance of selecting the PS model collaboratively, we designed quasi-experiments based on a real electronic healthcare database, where only the potential outcomes were manually generated, and the treatment and baseline covariates remained unchanged. Results showed that the C-TMLE algorithm outperformed other competing estimators for both point estimation and confidence interval coverage. In addition, the PS model selected by C-TMLE could be applied to other PS-based estimators, which also resulted in substantive improvement for both point estimation and confidence interval coverage. We illustrate the discussed concepts through an empirical example comparing the effects of non-selective nonsteroidal anti-inflammatory drugs with selective COX-2 inhibitors on gastrointestinal complications in a population of Medicare beneficiaries.

연구 동기 및 목표

  • 고차원 데이터에서 성향스코어 모델 선택에 있어 외부 교차검증의 한계를 해결하기 위해.
  • 치료 예측에만 집중하기보다 인과적 효과 추정의 목표 인과파rameter를 모델 선택에 통합함으로써 인과추론을 향상시키기 위해.
  • C-TMLE 기반 LASSO 모델 선택이 고차원 환경에서 점추정 및 신뢰구간 성능을 향상시키는지 평가하기 위해.
  • 협업적으로 선택된 모델이 다른 PS 기반 추정기로 이식 가능한지 평가하기 위해.
  • 실제 전자 건강기록 데이터와 준실험적 시뮬레이션을 사용하여 결과를 검증하기 위해.

제안 방법

  • 고차원 성향스코어 모델에서 LASSO의 조정파ram터를 선택하기 위해 C-TMLE를 활용하는 협업 제어 LASSO 프레임워크를 제안한다.
  • 치료 효과 추정에서 편향과 분산을 균형 있게 조절하기 위해 협업 학습을 통한 대상 최소 손실 기반 추정(TMLE)을 사용한다.
  • 초기 모델 피팅에는 외부 교차검증을 사용하지만, 인과파ram터에 초점을 맞춰 C-TMLE를 적용하여 모델 선택을 정밀화한다.
  • 두 개의 초기 추정기(단순 모델 및 기본 및 고차원 성향스코어(hdPS) 공변량을 포함한 슈퍼러닝 모델)를 사용하여 C-TMLE 알고리즘을 적용한다.
  • 교차검증된 이항 분산과 추정 정확도 및 신뢰구간 커버리지의 추정기 간 비교를 통해 모델 성능을 평가한다.
  • 선택된 성향스코어 모델을 여러 PS 기반 추정기(예: IPW, AIPW)에 적용하여 협업 선택의 일반화 가능성 테스트를 수행한다.

실험 결과

연구 질문

  • RQ1C-TMLE를 통한 협업 모델 선택이 고차원 환경에서 외부 교차검증 대비 평균 치료 효과의 점추정 성능을 향상시키는가?
  • RQ2C-TMLE 기반 LASSO 모델 선택은 인과효과 추정기의 신뢰구간 커버리지 및 길이에 어떤 영향을 미치는가?
  • RQ3C-TMLE로 선택된 성향스코어 모델을 비협업 기반 성향스코어 추정기에 효과적으로 이식할 수 있는가?
  • RQ4다른 초기 추정기(단순 대비 슈퍼러닝)가 C-TMLE에서 최종 모델 선택 및 추정 정확도에 어떤 영향을 미치는가?
  • RQ5인과추론에서 교란요인 통제를 목적으로 할 때, 외부 교차검증은 성향스코어 모델 선택에 있어 최적의 방법이 아닌가?

주요 결과

  • C-TMLE1 및 C-TMLE0 추정기는 시뮬레이션에서 점추정 정확도와 신뢰구간 커버리지 측면에서 최고의 성능을 기록했다.
  • C-TMLE로 선택된 모델은 λ = 0.000238로 총 166개의 공변량을 포함하였으며, 특히 예측력은 약한데도 강한 교란요인 효과를 지닌 hdPS 변수들을 더 많이 선택하였다.
  • 다른 PS 기반 추정기(예: IPW)에 적용했을 때, 협업적으로 선택된 모델은 점추정 성능을 크게 향상시켰다. 유일한 예외는 보편적인 IPW 추정기였다.
  • 단순 초기 추정기를 사용한 C-TMLE1 추정기는 교차검증 이항 분산 1.199632를 기록했으며, CV.LASSO의 1.199288보다 略로 높았지만, 인과추론 성능 측면에선 더 우수하였다.
  • 실증 분석 결과, COX-2 억제제 대비 비선택적 NSAIDs에 대한 평균 덧셈 치료 효과는 -0.249%로 추정되었으나 통계적으로 유의미하지 않았다.
  • 연구는 외부 교차검증이 최적의 교란요인 통제를 위해 부적합하며, C-TMLE 기반 앙상블 학습이 성향스코어 모델 선택에 있어 더 나은 대안임을 결론 내렸다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.