Skip to main content
QUICK REVIEW

[論文レビュー] Collaborative-controlled LASSO for Constructing Propensity Score-based Estimators in High-Dimensional Data

Cheng Ju, Richard Wyss|arXiv (Cornell University)|Jun 30, 2017
Advanced Causal Inference Techniques参考文献 28被引用数 9
ひとこと要約

本稿は、高次元データにおける傾向スコア推定のための共同制御付きLASSO手法を提案し、C-TMLEを用いて治療予測と因果効果推定を同時に最適化する。C-TMLEを用いたモデル選択が、平均治療効果推定のバイアス低減と信頼区間カバレッジの向上において、従来の外部交差検証を上回ることを示している。

ABSTRACT

Propensity score (PS) based estimators are increasingly used for causal inference in observational studies. However, model selection for PS estimation in high-dimensional data has received little attention. In these settings, PS models have traditionally been selected based on the goodness-of-fit for the treatment mechanism itself, without consideration of the causal parameter of interest. Collaborative minimum loss-based estimation (C-TMLE) is a novel methodology for causal inference that takes into account information on the causal parameter of interest when selecting a PS model. This "collaborative learning" considers variable associations with both treatment and outcome when selecting a PS model in order to minimize a bias-variance trade off in the estimated treatment effect. In this study, we introduce a novel approach for collaborative model selection when using the LASSO estimator for PS estimation in high-dimensional covariate settings. To demonstrate the importance of selecting the PS model collaboratively, we designed quasi-experiments based on a real electronic healthcare database, where only the potential outcomes were manually generated, and the treatment and baseline covariates remained unchanged. Results showed that the C-TMLE algorithm outperformed other competing estimators for both point estimation and confidence interval coverage. In addition, the PS model selected by C-TMLE could be applied to other PS-based estimators, which also resulted in substantive improvement for both point estimation and confidence interval coverage. We illustrate the discussed concepts through an empirical example comparing the effects of non-selective nonsteroidal anti-inflammatory drugs with selective COX-2 inhibitors on gastrointestinal complications in a population of Medicare beneficiaries.

研究の動機と目的

  • 高次元データにおける傾向スコアモデルの選択における外部交差検証の限界を解決すること。
  • 治療予測にのみ注目するのではなく、因果効果推定の目的変数をモデル選択に統合することで、因果推論を改善すること。
  • C-TMLEベースのLASSOモデル選択が、高次元設定における点推定および信頼区間の性能を向上させるかどうかを評価すること。
  • 共同的に選択されたモデルが他のPSベースの推定器へも適用可能かどうかを評価すること。
  • 実際の電子医療データと準実験的シミュレーションを用いて、主な発見を検証すること。

提案手法

  • C-TMLEを用いてLASSOのチューニングパラメータを高次元傾向スコアモデルで選択する共同制御付きLASSOフレームワークを提案する。
  • 治療効果推定におけるバイアスと分散のバランスをとるために、共同学習を用いたターゲット最小損失ベース推定(TMLE)を採用する。
  • 初期モデル適合には外部交差検証を用いるが、因果パラメータの目的に応じてC-TMLEを用いてモデル選択を精緻化する。
  • C-TMLEアルゴリズムに、ナイーブモデルとベースラインおよび高次元傾向スコア(hdPS)共変量を含むスーパー・ラーナー・モデルの2つの初期推定器を適用する。
  • 交差検証済み二項devianceを用いてモデル性能を評価し、推定精度および信頼区間カバレッジを異なる推定器間で比較する。
  • 選択されたPSモデルを、IPWやAIPWなどの複数のPSベース推定器に適用し、共同選択の一般化可能性を検証する。

実験結果

リサーチクエスチョン

  • RQ1C-TMLEによる共同モデル選択は、高次元設定において外部交差検証に比べ、平均治療効果の点推定を改善するか?
  • RQ2C-TMLEベースのLASSOモデル選択は、因果効果推定器の信頼区間カバレッジと長さにどのように影響を与えるか?
  • RQ3C-TMLEで選択されたPSモデルを非共同的PS推定器に効果的に転用できるか?
  • RQ4初期推定器(ナイーブ対スーパー・ラーナー)の違いが、C-TMLEにおける最終的なモデル選択および推定精度にどのように影響するか?
  • RQ5因果推論における交絡要因の制御を目的とする場合、外部交差検証はPSモデル選択において不適切であるか?

主な発見

  • C-TMLE1およびC-TMLE0推定器は、シミュレーションにおいて点推定精度と信頼区間カバレッジの両面で最良の性能を示した。
  • C-TMLEで選択されたモデルは166個の共変量を含み、λ = 0.000238であった。これは、予測力が弱くても交絡効果が強いhdPS変数を多く含むため、CV.LASSOよりも多くの交絡要因を特定していた。
  • 他のPSベース推定器に適用した場合、共同的に選択されたモデルは点推定を著しく改善したが、バニラIPW推定器を除いては同様であった。
  • ナイーブ初期推定器を用いたC-TMLE1推定器は、交差検証済み二項devianceが1.199632とやや高かった(CV.LASSOの1.199288よりわずかに高いが)、因果推論性能は優れていた。
  • 実データ分析では、COX-2阻害薬と非選択的NSAIDsの平均加法的治療効果は-0.249%であったが、有意ではなかった。
  • 本研究は、外部交差検証は最適な交絡要因の制御に不十分であり、C-TMLEベースのアンサンブル学習がPSモデル選択のより優れた代替手段であると結論づけた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。